Two weeks ago Alibaba previewed Qwen3.8-Max with a claim and nothing else. No benchmark table, no model card, no license, no per-token price. We called that a press release rather than a model, and said the gap between announcement and artifact was becoming a credibility problem for Chinese labs generally.
Alibaba closed it. Qwen3.8-Max went generally available on August 3 with a full benchmark table, standard pricing, disclosed active parameters, and a commitment to publish open weights the week of August 10. That last item is the one that matters most, because no Max-tier Qwen has ever shipped with open weights. If it lands, it will be the first self-hostable flagship in the family’s history.
What the Numbers Actually Say
The model is a 2.4-trillion-parameter mixture of experts activating 95 billion per token, multimodal across text, images, and video, with a million-token context window. The practical ceilings are 991K tokens of input, 983K with thinking enabled, and 131K of output, with reasoning chains running to 262K.
| Benchmark | Qwen3.8-Max | Comparison |
|---|---|---|
| Terminal-Bench 2.1 | 86.6 | GPT-5.6 Sol 88.8; Fable 5 84.6; Opus 4.8 84.6 |
| SWE-bench Pro | 67.7 | Fable 5 80.0 |
| FrontierSWE | 73.5 | Fable 5 88.8 |
| DeepSWE 1.1 | 56.6 | Qwen3.7-Max 21.6 |
| GPQA Diamond | 92.6 | Qwen3.7-Max 92.4 |
| PaperBench | 93.0 | — |
| OSWorld-Verified | 86.1 | — |
| OmniDocBench 1.5 | 92.1 | — |
| Parametric CAD Bench | 91.5 | — |
| IFBench | 82.8 | — |
| JobBench | 53.4 | Qwen3.7-Max 31.3 |
Now that the table exists, the July claim can be checked, and it was selectively true. Alibaba’s team said the model was “second only to Fable 5” among systems it tested. On Terminal-Bench it actually beats Fable 5 by two points and loses to GPT-5.6 Sol. On the two hardest software engineering evaluations it loses to Fable 5 badly, trailing by 12.3 points on SWE-bench Pro and 15.3 on FrontierSWE. A model can be ahead of Anthropic on agentic terminal work and well behind it on repository-scale code, and Qwen3.8-Max is both.
The more informative column is the one comparing against Qwen’s own previous flagship. GPQA Diamond moved 0.2 points, which is saturation. DeepSWE 1.1 moved from 21.6 to 56.6 and JobBench from 31.3 to 53.4. That is the same shape DeepSeek just published for V4-Flash: static academic reasoning scores, large agent-task gains. Two Chinese labs in one week reporting that their improvements landed almost entirely in tool use and long-horizon work is a pattern rather than a coincidence.
Pricing is $2.00 per million input tokens and $6.00 output, with implicit cache reads at $0.25, explicit cache writes at $2.50, and explicit cache reads at $0.17. That undercuts Kimi K3’s $3.00 and $15.00 while sitting well above DeepSeek’s $0.14 and $0.28. Alibaba shares rose about 7% in Hong Kong on the announcement.
The Open-Weight Commitment Is the Story
Qwen has open-weighted plenty of models. It has never open-weighted a Max. The Max tier has been the closed commercial flagship in every prior generation, which is exactly the structure OpenAI and Anthropic use and exactly the structure the open-model argument exists to pressure. Alibaba says weights for both Qwen3.8-Max and a smaller Qwen3.8-27B will publish the week of August 10 on Hugging Face and ModelScope.
The 27B is the one most teams will actually use. A dense model in that class quantizes down to something that runs on a single consumer GPU, which puts it in the category Kimi K3’s 1.4 terabytes will never reach. Alibaba pairing a 2.4T flagship with a 27B sibling in the same release is a deliberate answer to the deployability problem the large Chinese open models have created for themselves this month, and it is a better answer than shipping the flagship alone.
Whether the flagship weights matter practically is a separate question, and they will matter to very few, because a 2.4T MoE carries the same procurement problem as K3. What the release does is set a precedent about what a Chinese lab is willing to give away at the top of its lineup, and that precedent applies pressure to every lab that isn’t.
The License Is Still Blank
One thing from the July critique has not been resolved. Alibaba has announced no license. Qwen’s open releases have historically been Apache 2.0, which is why parts of the community are assuming something similarly permissive, but assumption is not disclosure and a Max-tier flagship is exactly the release where a lab would introduce revenue thresholds if it were going to. Moonshot did precisely that with K3, moving from a standard permissive license to a bespoke one with commercial gates once the model got large enough to matter.
So the scorecard is better than it was on July 22 without being complete. The benchmarks, the pricing, and the active parameter count are all published now. The weights have a date but not a delivery. The license is still unknown, and it happens to be the piece that determines what the weights are actually worth when they arrive.
There is also a governance dimension worth flagging for anyone considering this in an enterprise setting. Running Qwen through Alibaba’s hosted endpoints puts inference on infrastructure subject to Chinese state law, which is a compliance question independent of the model’s quality and one that self-hosting is the only real answer to. That is the strongest practical argument for caring about the open-weight release regardless of what the flagship costs to run.
The reasonable position for now is to treat Qwen3.8-Max as a credible model with verified numbers and an unverified license, evaluate it on the workloads where it actually leads rather than on the headline claim, and wait a week before deciding what the open-weight announcement is worth. Alibaba has spent the last two weeks earning back the benefit of the doubt it burned in July, which is more than most labs bother to do after shipping an adjective. Whether that extends to the license is the one thing still worth withholding judgment on.
Sources
- Alibaba Qwen Releases Qwen3.8-Max — MarkTechPost
- Qwen 3.8 Max Ships: 2.4T MoE, 1M Context, Open Weights Next Week — Developers Digest
- Alibaba says Qwen3.8-Max tops Kimi K3 on some benchmarks — Bloomberg via Techmeme
- Alibaba unveils 2.4-trillion-parameter Qwen3.8-Max — Tech Startups
