📊 Full opportunity report: Is Qwen3.8-Max The New Second In AI? The Data Might Look Different on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Alibaba announced the broad availability of Qwen3.8-Max, a 2.4 trillion-parameter AI model with confirmed benchmark results. The open weights will be released soon, marking a significant step in large-scale open AI models. The claim of being second only to Fable 5 is based on selective benchmarks, and full details are still emerging.
Alibaba has officially launched Qwen3.8-Max, confirming it as a 2.4 trillion-parameter AI model with strong benchmark performance. The company revealed detailed benchmark results and announced the upcoming release of open weights, marking a significant milestone in large-scale open AI models.
On August 3, Alibaba published the full benchmark table for Qwen3.8-Max, previously previewed in July. The model features approximately 95 billion active parameters with a total of 2.4 trillion parameters, built on the Qwen3.5 architecture and employing sparse mixture-of-experts technology. It is multimodal, capable of processing text, images, and videos, with text output.
The benchmark results place Qwen3.8-Max at the top of several tests, including Terminal-Bench 2.1 at 86.6, surpassing models like Claude Opus 4.8 and Claude Fable 5, but trailing behind GPT-5.6 Sol at 88.8. It also leads in PaperBench (93.0) and performs well in multimodal and agentic tasks, such as OSWorld-Verified (86.1) and Parametric CAD Bench (91.5). However, it trails significantly on deep software engineering benchmarks like SWE-bench Pro (67.7) and FrontierSWE (73.5).
Alibaba emphasized that the model has demonstrated substantial improvements in agentic tasks, especially in long-horizon reasoning, achieving new benchmarks compared to its predecessor, Qwen3.7-Max.
For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.
▲ All performance figures: Alibaba’s own harnessThe claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.
“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.
“Qwen3.8 is going open-weight” describes three things with very different deployment realities.
OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.
A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.
The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.
Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.
- The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
- More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
- If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
- The 27B sibling could become the best local agent model on hardware people already own.
- Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
- The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
- “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
- Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
and it says “second only” depends entirely on which row you read.
Implications of Alibaba's Benchmark and Open-Weight Release
The announcement confirms Alibaba's position as a major player in large-scale AI development, with the release of a 2.4 trillion-parameter model and detailed benchmark data. The open weights scheduled for next week will enable wider access and experimentation, potentially reshaping competition in the AI ecosystem. However, the model's selective benchmark performance indicates that claims of being second only to Fable 5 are based on specific tests, not a comprehensive ranking.
This development matters because it signals increased transparency and availability of large-scale models, fostering innovation and competition. The performance in agentic and multimodal tasks highlights the model's practical capabilities, though limitations in software engineering benchmarks suggest areas for further improvement.

AI Systems Performance Engineering: Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorch
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background on Alibaba's AI Model Development and Strategy
Alibaba's AI journey has been marked by stealth previews and strategic releases, notably the July preview of Qwen3.8-Max, which was initially presented as a slogan claiming second place behind Fable 5. The model's parameters and capabilities were kept under wraps until the full benchmark table was published on August 3. Prior to this, Alibaba's models, such as Kimi K3 and earlier Qwen versions, had demonstrated rapid progress in multimodal and agentic tasks, often through scaling and reinforcement learning techniques.
The company's approach has involved staged disclosures, with the recent benchmarks providing transparency into the model's strengths and weaknesses, particularly in agentic reasoning and software engineering benchmarks. The upcoming open weights are part of Alibaba's broader strategy to foster open AI development and challenge existing leaders like OpenAI and Anthropic.
"Qwen3.8-Max represents a significant leap in our multimodal and agentic capabilities, with benchmark results confirming its competitive edge."
— Alibaba spokesperson
Remaining Questions About Model Capabilities and Release Details
It is not yet clear whether the open weights will match the full 2.4 trillion-parameter checkpoint or be a distilled version suitable for single-machine inference. The licensing terms are still unpublished, which could influence adoption and use. Additionally, the model's performance on software engineering benchmarks remains significantly lower than top models, raising questions about its overall versatility and readiness for deployment in certain domains.
Further details on the licensing, the exact size of the open weights, and real-world performance in diverse applications are still forthcoming.
Next Steps for Alibaba's AI Model Deployment and Community Access
Alibaba will release the open weights next week, enabling developers and researchers to evaluate and deploy the model independently. The company may also publish additional benchmarks and documentation to clarify licensing and technical specifications. The broader AI community will closely monitor the model's performance in real-world tasks, especially in agentic and software engineering applications.
Expect further updates from Alibaba regarding licensing terms, model capabilities, and potential integrations into commercial and research workflows.
Key Questions
What is Qwen3.8-Max and why is it significant?
Qwen3.8-Max is Alibaba's newly announced 2.4 trillion-parameter multimodal AI model, with benchmark results indicating strong performance, especially in multimodal and agentic tasks. Its upcoming open weights could impact AI development broadly.
How does Qwen3.8-Max compare to other models like Fable 5?
Based on Alibaba's benchmark data, Qwen3.8-Max appears to be second only to Fable 5 in certain tests, but this is based on selective benchmarks. Its performance in software engineering benchmarks is significantly lower.
When will the open weights be available?
The open weights for Qwen3.8-Max are scheduled to be released next week, enabling broader access for developers and researchers.
What are the limitations of Qwen3.8-Max?
The model trails behind in software engineering benchmarks and its licensing terms remain unpublished, which could affect how it is used and integrated.
What does this mean for the AI industry?
This development signifies increased transparency and openness in large-scale AI models, potentially intensifying competition and innovation across the industry.
Source: ThorstenMeyerAI.com