Back to Insights
AI

Kimi K3: Moonshot Ships 2.8 Trillion Parameters — and Its Own Chart Shows It Losing

Moonshot's Kimi K3 ships 2.8T parameters and 1M context. Artificial Analysis ranks it #4 — but the weights and the license don't exist yet.

S5 Labs Team July 16, 2026

Moonshot AI launched Kimi K3 today across Kimi.com, Kimi Work, Kimi Code and the Kimi API, and it is billable right now — OpenRouter has it live, served solely by Moonshot, at $3.00 per million input tokens and $15.00 per million output. The official blog puts it at 2.8 trillion total parameters with a 1,048,576-token context window, making it the largest model any lab has ever announced as open-weight, roughly 1.75x DeepSeek V4-Pro’s 1.6T. The weights, Moonshot says, will be released by July 27.

That is the launch. The more interesting document is the benchmark table Moonshot published next to it, because on Moonshot’s own numbers, K3 loses.

Scorecard of Kimi K3's three launch claims: capability is real (Artificial Analysis Intelligence Index 57, #4 of 189, up from K2.6's 44); candor is real (Moonshot's own chart shows K3 losing to Claude Fable 5 on most rows); but openness is a date, not a fact — no repo, model card, or license as of July 16, with weights only promised by July 27. Specs: 2.8T parameters, 1M context, 3.00/15.00 per million tokens.

What Shipped

K3 breaks from the K2 line rather than incrementing on it. K2.6 was a one-trillion-parameter mixture of experts with 32 billion active per token, 384 experts and a 256K context; K3 runs Stable LatentMoE with 896 experts, 16 activated per token, and a 1M-token context, on MXFP4 weights trained quantization-aware from the SFT stage on. Attention is Kimi Delta Attention plus Attention Residuals — not a launch-week coinage, KDA comes out of the Kimi Linear paper from October 2025. That is 2.8x the parameters and 4x the context of a model that shipped three months ago, with K2.7 Code in between. Moonshot does not publish an active-parameter count for K3, though it disclosed 32B of 1T for K2.6 without hesitation, and sparse models are sold on exactly that ratio.

The parameter count is itself a small lesson. Aggregators spent the run-up publishing a leaked 2.5T figure as fact; TechCrunch, citing Financial Times sources, hedged correctly to “between 2 trillion and 3 trillion.” The lab’s own number was available within hours and much of the search-visible coverage still carries the wrong one — the same failure mode that invented a launch date for Gemini 3.5 Pro.

Moonshot’s Own Chart Says It Loses

This is the part worth crediting. The blog says K3’s overall performance “still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol,” and the limitations section concedes “a noticeable gap in user experience” plus “excessive proactiveness” — the model taking autonomous decisions nobody asked it for. On the comparison table Moonshot itself published, K3 loses to Claude Fable 5 on the clear majority of rows and trails GPT-5.6 Sol on many of the rest. Several outlets covering the launch are louder about K3 than Moonshot is. Set that against the K2.7 Code launch six weeks ago, where every headline benchmark was an in-house suite and every chart flattered the model. A lab shipping a table it visibly loses on is rare enough to name plainly.

The wins cluster in agentic and retrieval-shaped work rather than raw generality: BrowseComp at 91.2, Terminal Bench 2.1 at 88.3, OmniDocBench at 91.1, DeepSearchQA F1 at 95.0. There is also a DeepSWE figure, 67.5, from a lab that declined to submit K2.7 Code to that benchmark at all. All of it is self-reported at maximum reasoning effort — the lab grading its own exam.

The One Number That Isn’t Moonshot’s

Artificial Analysis has already scored K3 independently, and this is where the release earns its headline instead of asserting it. AA puts K3 at 57 on its Intelligence Index, ranked #4 of 189 models, up from 44 for K2.6 — thirteen points of independently measured gain in three months. That is the actual story here: the open-weight frontier now sits roughly one tier behind the closed frontier instead of two.

The same page carries the other half. K3 outputs 62.0 tokens per second, ranked #90 of 189, and AA’s own summary is that the model is “slower than average and very verbose” and “amongst the leading models in intelligence, but somewhat expensive,” at a blended $2.31 per million tokens. TestingCatalog, evaluating an Arena stealth model widely believed to be a K3 preview, clocked one task at 35 minutes and called it “VERY slow, even slower than Fable.” And reasoning_effort is locked to max, with more levels “coming soon,” so there is no lever to turn the burn down.

Those halves belong in one sentence. The intelligence is independently confirmed; the cost advantage is not, because a verbose model at 40% off per token that emits three times as many tokens costs more per finished job. Measure total spend per completed task, not price per token.

The Price Isn’t the Cheap-Chinese-Model Story

At $3.00 in and $15.00 out, K3 is exactly 40% below Claude Opus 4.8 at $5.00/$25.00 — the comparison the coverage reaches for, and the most flattering one available. Two things complicate it. Anthropic’s docs note that Opus 4.7 and later, Fable 5 and Sonnet 5 use a newer tokenizer producing roughly 30% more tokens for the same text, so the two models are not billing the same unit; 40% cheaper per token is not 40% cheaper per unit of work. And $3.00/$15.00 is precisely what Claude Sonnet 5 costs from September 1, once its introductory pricing lapses — against Anthropic’s mid-tier, the realistic comparison for most buyers, the discount is zero.

It is also roughly 3x what Moonshot charges for its own K2.6 at $0.95/$4.00, a model scoring a respectable 44 on the same index whose weights you can actually download today. What rescues K3’s economics is Moonshot’s claim of a cache-hit rate above 90% on coding workloads, which would put repeated-context work near the $0.30 cache-hit rate rather than $3.00 — a vendor claim about your workload rather than a measurement of it, and the first thing to test rather than believe.

”Open” Is a Deadline, Not a License

Here is the hole, and it is not a footnote. As of today there is no K3 repository on Hugging Face, no model card, no LICENSE file, and no statement from Moonshot naming a license at all. OpenRouter labels it “open-weight” and specifies nothing further. The Modified MIT terms everyone assumes are an inference from the K2 lineage: K2.6’s license reads as plain MIT until you cross 100 million monthly active users or $20 million in monthly revenue, at which point you must display the model’s name in your product — but that text names the specific string “Kimi K2.6,” so K3’s terms are genuinely unknown until someone publishes them. “The world’s first open 3T-class model” is Moonshot’s phrase, rounding 2.8T up, and it deserves attribution rather than repetition.

Even if July 27 lands as promised, “open” here does not mean what a reader hoping to run it locally wants it to mean. At 2.8T parameters in MXFP4 the weights come to roughly 1.4 terabytes, and Hacker News commenters put self-hosting north of $100,000 in GPUs — in a market where hardware prices are moving the wrong way. The practical path for essentially everyone is the API or a third-party host, operationally identical to using a closed model. Open weights at this scale buy ecosystem leverage and insurance against lock-in, not a box under your desk.

What July 27 Actually Decides

Two of the three claims here hold up under an outside look. The capability is real and independently measured — #4 of 189 on an index Moonshot does not run is a real result at a new scale — and the candor about where it still loses is worth crediting from the lab that shipped K2.7 Code on nothing but its own suites. The third claim carries the headline, and it is a date. Until a repository and a license file exist, “largest open model in the world” describes a commitment eleven days out rather than a thing that happened. Pilot it on the API against a real workload and measure what a finished task costs; do not architect around downloadable weights until they are downloadable. Whether this was a landmark or a press cycle is a question with a scheduled answer.

Key Details

SpecDetail
LabMoonshot AI
ModelKimi K3
Total Parameters2.8 trillion (Stable LatentMoE)
Experts896 total, 16 activated per token
Active ParametersNot disclosed
Context Window1,048,576 tokens (1M)
AttentionKimi Delta Attention + Attention Residuals
PrecisionMXFP4 weights, MXFP8 activations (QAT from SFT onward)
LicenseNot stated by Moonshot (K2.6 shipped under Modified MIT)
WeightsPromised by July 27, 2026 — not published as of today
Input Pricing$3.00 / 1M tokens ($0.30 cache-hit)
Output Pricing$15.00 / 1M tokens
Reasoning EffortLocked to max (more levels “coming soon”)
Artificial AnalysisIntelligence Index 57 (#4 of 189); 62.0 tok/s (#90 of 189)
AvailabilityKimi API, OpenRouter

Update — July 27: Moonshot published the weights on schedule, which settles the open question this article raised. The Hugging Face repository is 118 files totaling about 1.56TB, under a bespoke MIT-derived “Kimi K3 License” that adds a separate-agreement requirement for model-as-a-service operators above $20M in group revenue and an attribution requirement above 100M monthly active users. The active-parameter count Moonshot withheld at launch is now disclosed at 104.2 billion of 2.78 trillion. The practical constraint turned out to be hardware rather than licensing: roughly 1.4TB resident at MXFP4 before KV cache, which means eight B300s to serve it on a single node.

Sources

Want to discuss this topic?

We'd love to hear about your specific challenges and how we might help.