Back to Insights
AI

Alibaba Answered Kimi K3 in Three Days With a 2.4-Trillion-Parameter Model — and Zero Benchmarks

Alibaba's Qwen3.8-Max claims to trail only Fable 5 — with no benchmarks, license, or open weights. China's open-model race has a credibility problem.

S5 Labs Team July 22, 2026

On July 19, at the World AI Conference in Shanghai, Alibaba previewed Qwen3.8-Max — a 2.4-trillion-parameter multimodal model its own team called “second only to Fable 5” among the systems it tested against. It shipped that claim with no benchmark table, no model card, no license, no independent evaluation, and no open weights. Three days earlier, Moonshot AI had launched Kimi K3: 2.8 trillion parameters, open weights on a dated schedule, published benchmarks, published pricing, and an independent Artificial Analysis score to back the headline.

Put those two releases next to each other and the shape of the week is obvious. One lab shipped a model. The other shipped a press release timed to sit on top of it.

Qwen3.8-Max versus Kimi K3, a claim next to a model. Alibaba's Qwen3.8-Max, marked as a claim: 2.4 trillion parameters, previewed July 19 at the World AI Conference, said to be "second only to Fable 5" — but shipped with no benchmark table, no model card, no license, and no open weights, promised only "soon." It spent an adjective. Moonshot's Kimi K3, marked as a model: 2.8 trillion parameters, open weights committed for July 27, published benchmarks, and an independent Artificial Analysis Intelligence Index score of 57, ranked #4 of 189. It spent a date. The twist along the bottom: Alibaba owns roughly 36% of Moonshot through an $800 million 2024 investment, so it is racing a rival it helped fund; Alibaba's US shares rose more than 3% on the preview; and China's Ministry of Commerce is weighing restrictions on overseas downloads of exactly these open-weight models.

What Alibaba Actually Previewed

Qwen3.8-Max is real in the sense that you can pay to use it today. It is live as Qwen3.8-Max-Preview through Alibaba’s Token Plan, Qoder, and QoderWork, running at roughly 10% of eventual standard pricing during the preview — a credit-based ladder from a $6 Lite tier to a $68 Pro tier that allows six to eight concurrent agents. The architecture is a sparse mixture-of-experts, multimodal across text, images, video, and documents, with a one-million-token context window. As a product preview, it exists.

As a model launch, it doesn’t yet. What’s missing is everything that would let anyone check the headline. Alibaba’s Max-tier releases normally ship with the weights available and numbers attached; this one inverted the pattern. There is no active-parameter count, which is the number that actually determines inference cost on a sparse model — 2.4 trillion total tells you almost nothing about what a query costs to serve. There is no entry on Artificial Analysis, LMArena, or the Hugging Face leaderboard. Open weights are promised “soon,” with no date and no named license. NotebookCheck, eWeek, and Quartz all led with the same observation: the claim arrived naked.

The “Checkpoint Gap”

The independent analyst Julien Simon gave this maneuver a name worth borrowing: the checkpoint gap. A lab announces a model — a training checkpoint — the moment a rival ships, not the moment its own model is ready to be judged, because the goal is to occupy the narrative slot rather than to win the benchmark. “In the checkpoint gap,” Simon wrote, “the only hard currency is a date, and Alibaba isn’t spending any.” Kimi K3 spent one: its weights are committed to Hugging Face for July 27, on a public clock everyone can watch. Qwen3.8-Max spent an adjective.

This is the same failure mode we flagged with Kimi K3 itself, from the other direction. K3’s capability was independently confirmed while its openness was still just a promised date — a real model with one unshipped claim. Qwen3.8-Max is further back: the capability claim is the unshipped part. The discipline for reading either is the same one that applies to any AI benchmark — a number a lab reports about its own model, on a test it chose, at a setting it picked, is a marketing artifact until someone outside the building reproduces it. Qwen3.8-Max hasn’t reached even that first bar, because there is no number to reproduce.

The Part Nobody Says Out Loud

The rush makes more sense once you know who owns whom. Alibaba owns roughly 36% of Moonshot AI — the maker of Kimi K3 — through an investment of about $800 million in early 2024, when Moonshot was valued near $2.5 billion. Investor Michael Burry called Qwen3.8-Max a “clear response” to K3, and it plainly is. But it is a response to a rival Alibaba helped fund. The company that shipped a naked capability claim to blunt Kimi K3’s launch is the same company that owns better than a third of the lab it’s racing.

Alibaba’s US-listed shares rose more than 3% on the news, to around $119. That is the actual currency being spent here, and it clears immediately — a preview at a conference moves a stock in a way that a fair benchmark comparison three weeks later never will. The market rewarded the announcement, not the model, which is precisely why the announcement came before the model.

Why the Race Is Getting Sloppy

Step back from the two labs and the context explains the sloppiness. Over the past year, China’s open-weight models became the default for cost-conscious developers worldwide. DeepSeek’s cost disruption cracked the door; Qwen, Kimi, and Zhipu’s GLM line walked through it, precisely because their weights are downloadable, fine-tunable, and self-hostable in a way the API-only Western leaders are not. When a category is winning globally and four domestic labs are fighting for the same top slot, the temptation to announce before you’re ready gets stronger, not weaker. Qwen3.8-Max is what that pressure looks like when it wins.

Worth noting what did not ship this week, because absence is data too: no new DeepSeek. R2 remains unreleased amid reports its founder is unhappy with it, and V4-Pro’s move from preview to general availability stayed unconfirmed. The lab that started this whole cycle sat the week out while its rivals traded headlines.

Beijing May Close the Tap It Can’t See

The genuine twist isn’t competitive — it’s regulatory, and it cuts against the whole open-weight story. Since early July, China’s Ministry of Commerce has reportedly been in talks with Alibaba, ByteDance, and Zhipu about restricting overseas access to advanced and unreleased Chinese models. Reuters first reported the discussions on July 7; the Financial Times added specifics around July 21, including possible limits on foreign downloads of model weights and training-data export, and — further out — curbs that could bar Qualcomm and TSMC from manufacturing advanced chips based on Chinese firms’ designs. API and cloud access would reportedly stay open; it’s the downloadable weights, the thing that made these models the global default, that are on the table.

The irony arrived on schedule and from an unlikely source. In the same week, Hugging Face disclosed that it fought off an autonomous AI cyberattack partly by running forensics on GLM-5.2, Zhipu’s open Chinese model, after commercial Western models refused to examine the attack code. A Western AI company reached for a Chinese open-weight model in a crisis at the exact moment Beijing began weighing whether to restrict foreign access to it. Even David Sacks, the former White House AI czar, argued against the mirror-image US instinct: “There’s no reason to limit American models on tasks that Chinese models handle without issue.” The open-weight flow that both governments are eyeing as a lever is, in practice, load-bearing infrastructure for the other side’s defenders.

What This Means If You Actually Build With These Models

Nothing here demands action today, and one thing specifically warns against it. Qwen3.8-Max-Preview is not something to evaluate, budget for, or mention to a client yet. There is no model card to read, no license to clear with your legal team, no independent score to weigh, and no committed date for the weights. Treat it as a stock catalyst that happens to have a parameter count, and revisit it the day Alibaba publishes something Artificial Analysis or LMArena can actually score. Judge it on total cost per finished task once real numbers exist — a headline parameter count has never once predicted what a workload actually costs to run.

If your stack already runs an open Chinese model — Qwen, DeepSeek, Kimi, GLM — via API or self-hosted for cost reasons, the Ministry of Commerce talks are a calendar note, not a re-architecture order. The proposal is at the consultation stage, appears aimed at future and frontier releases and at weight downloads specifically, and reportedly leaves API and cloud access intact. The prudent hedge is the one that was always prudent with any single-vendor dependency: know which model you’re locked to, keep a tested fallback on a different provider, and don’t let a self-hosted weight file become a thing you can’t replace if the download tap closes. That was good practice before Beijing said a word. This week just gave it a date to point at.

Key Details

ItemQwen3.8-Max-PreviewKimi K3
LabAlibabaMoonshot AI
AnnouncedJuly 19, 2026 (WAIC Shanghai)~July 16, 2026
Total parameters2.4 trillion (sparse MoE)2.8 trillion (MoE)
Context window1,000,000 tokens1,048,576 tokens
Open weightsPromised “soon” — no date, no licenseCommitted July 27, 2026
Published benchmarksNoneSelf-reported + independent (AA)
Independent scoreNoneArtificial Analysis Index 57 (#4 of 189)
AvailabilityToken Plan, Qoder, QoderWork (~10% preview pricing)Kimi API, OpenRouter
The tellA claim without a numberA model with one unshipped promise

Update — August 3: Alibaba closed most of the gap this article described. Qwen3.8-Max went generally available with a full benchmark table, standard pricing at $2.00 input and $6.00 output per million tokens, a disclosed 95B active of 2.4T, and a commitment to publish open weights for both Qwen3.8-Max and a smaller Qwen3.8-27B the week of August 10. That would be the first open-weight Max-tier model in Qwen’s history. The published numbers also make the July claim checkable, and it was selective: Qwen3.8-Max beats Fable 5 on Terminal-Bench 2.1 (86.6 to 84.6) but trails it on SWE-bench Pro (67.7 to 80.0) and FrontierSWE (73.5 to 88.8). The license remains unannounced.

Sources

Want to discuss this topic?

We'd love to hear about your specific challenges and how we might help.