📊 Full opportunity report: Is Qwen3.8-Max The New Second In AI? The Data Might Look Different on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba announced the broad availability of Qwen3.8-Max, a 2.4 trillion-parameter AI model with confirmed benchmark results. The open weights will be released soon, marking a significant step in large-scale open AI models. The claim of being second only to Fable 5 is based on selective benchmarks, and full details are still emerging.

Alibaba has officially launched Qwen3.8-Max, confirming it as a 2.4 trillion-parameter AI model with strong benchmark performance. The company revealed detailed benchmark results and announced the upcoming release of open weights, marking a significant milestone in large-scale open AI models.

On August 3, Alibaba published the full benchmark table for Qwen3.8-Max, previously previewed in July. The model features approximately 95 billion active parameters with a total of 2.4 trillion parameters, built on the Qwen3.5 architecture and employing sparse mixture-of-experts technology. It is multimodal, capable of processing text, images, and videos, with text output.

The benchmark results place Qwen3.8-Max at the top of several tests, including Terminal-Bench 2.1 at 86.6, surpassing models like Claude Opus 4.8 and Claude Fable 5, but trailing behind GPT-5.6 Sol at 88.8. It also leads in PaperBench (93.0) and performs well in multimodal and agentic tasks, such as OSWorld-Verified (86.1) and Parametric CAD Bench (91.5). However, it trails significantly on deep software engineering benchmarks like SWE-bench Pro (67.7) and FrontierSWE (73.5).

Alibaba emphasized that the model has demonstrated substantial improvements in agentic tasks, especially in long-horizon reasoning, achieving new benchmarks compared to its predecessor, Qwen3.7-Max.

At a glance
updateWhen: announced August 3, 2023; benchmarks pu…
The developmentAlibaba officially announced and published benchmark data for Qwen3.8-Max, confirming it as a major new AI model with open weights forthcoming.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba's Benchmark and Open-Weight Release

The announcement confirms Alibaba's position as a major player in large-scale AI development, with the release of a 2.4 trillion-parameter model and detailed benchmark data. The open weights scheduled for next week will enable wider access and experimentation, potentially reshaping competition in the AI ecosystem. However, the model's selective benchmark performance indicates that claims of being second only to Fable 5 are based on specific tests, not a comprehensive ranking.

This development matters because it signals increased transparency and availability of large-scale models, fostering innovation and competition. The performance in agentic and multimodal tasks highlights the model's practical capabilities, though limitations in software engineering benchmarks suggest areas for further improvement.

AI Systems Performance Engineering: Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorch

AI Systems Performance Engineering: Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorch

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba's AI Model Development and Strategy

Alibaba's AI journey has been marked by stealth previews and strategic releases, notably the July preview of Qwen3.8-Max, which was initially presented as a slogan claiming second place behind Fable 5. The model's parameters and capabilities were kept under wraps until the full benchmark table was published on August 3. Prior to this, Alibaba's models, such as Kimi K3 and earlier Qwen versions, had demonstrated rapid progress in multimodal and agentic tasks, often through scaling and reinforcement learning techniques.

The company's approach has involved staged disclosures, with the recent benchmarks providing transparency into the model's strengths and weaknesses, particularly in agentic reasoning and software engineering benchmarks. The upcoming open weights are part of Alibaba's broader strategy to foster open AI development and challenge existing leaders like OpenAI and Anthropic.

"Qwen3.8-Max represents a significant leap in our multimodal and agentic capabilities, with benchmark results confirming its competitive edge."

— Alibaba spokesperson

Remaining Questions About Model Capabilities and Release Details

It is not yet clear whether the open weights will match the full 2.4 trillion-parameter checkpoint or be a distilled version suitable for single-machine inference. The licensing terms are still unpublished, which could influence adoption and use. Additionally, the model's performance on software engineering benchmarks remains significantly lower than top models, raising questions about its overall versatility and readiness for deployment in certain domains.

Further details on the licensing, the exact size of the open weights, and real-world performance in diverse applications are still forthcoming.

Next Steps for Alibaba's AI Model Deployment and Community Access

Alibaba will release the open weights next week, enabling developers and researchers to evaluate and deploy the model independently. The company may also publish additional benchmarks and documentation to clarify licensing and technical specifications. The broader AI community will closely monitor the model's performance in real-world tasks, especially in agentic and software engineering applications.

Expect further updates from Alibaba regarding licensing terms, model capabilities, and potential integrations into commercial and research workflows.

Key Questions

What is Qwen3.8-Max and why is it significant?

Qwen3.8-Max is Alibaba's newly announced 2.4 trillion-parameter multimodal AI model, with benchmark results indicating strong performance, especially in multimodal and agentic tasks. Its upcoming open weights could impact AI development broadly.

How does Qwen3.8-Max compare to other models like Fable 5?

Based on Alibaba's benchmark data, Qwen3.8-Max appears to be second only to Fable 5 in certain tests, but this is based on selective benchmarks. Its performance in software engineering benchmarks is significantly lower.

When will the open weights be available?

The open weights for Qwen3.8-Max are scheduled to be released next week, enabling broader access for developers and researchers.

What are the limitations of Qwen3.8-Max?

The model trails behind in software engineering benchmarks and its licensing terms remain unpublished, which could affect how it is used and integrated.

What does this mean for the AI industry?

This development signifies increased transparency and openness in large-scale AI models, potentially intensifying competition and innovation across the industry.

Source: ThorstenMeyerAI.com

You May Also Like

Green Screens Fail for One Reason—Here’s How to Fix It Fast

Ineffective green screen lighting causes failures; discover quick fixes to achieve seamless backgrounds and improve your filming results today.

USB Mic vs XLR Mic: The Decision Tree Every Creator Needs

If you’re choosing between a USB and XLR mic, consider your needs…

AR Displays: Waveguides, Micro‑OLED, and MicroLED

Keen to understand how waveguides, micro-OLED, and microLED transform AR displays and what this means for your immersive experience?

Data Centers Surges In Global Coverage

Data center mentions worldwide have increased 15 times according to GDELT, highlighting rapid expansion in infrastructure and digital connectivity.