Mistral launched a public preview of Mistral Large 4 on October 6, a mixture-of-experts model it describes as “1 trillion-parameter natively multimodal” with 52 billion active parameters, and said the weights will follow by the end of October. Mistral Today the model exists as an API on Mistral Studio with a published price, a docs page and a long benchmark post. The two things a self-hosting decision turns on, the checkpoint and its license, do not exist yet, and the independent record is one index score. The launch post names no license; the docs page tags the model “Open” where Mistral’s other models carry “Apache 2.0” or “Modified MIT”. Mistral docs, models
This is the same position Reflection’s Beam was in a day earlier, with one difference: Mistral has shipped large open checkpoints before, under more than one license. S5 Labs has not tested the model, and every benchmark below is Mistral’s unless marked as Artificial Analysis.
What you can use today
The preview is served through Mistral Studio and the standard endpoints: chat completions, conversations, agents and batch. The docs list structured outputs, function calling, document QnA and prefix completion, a 1M-token context window, and a 1.6B-parameter vision encoder. The model ID on the docs page is mistral-large-4 with one alias; the docs’ own compare links and OpenRouter’s listing use mistral-large-4-0. Mistral docs, Mistral Large 4 OpenRouter
Pricing per million tokens, from the docs page:
| List | Sale | |
|---|---|---|
| Input | $1.36 | $0.68 |
| Cached input | $0.14 | $0.07 |
| Output | $4.18 | $2.09 |
The sale is a straight 50% off with no end date on the page. Mistral docs, Mistral Large 4 For comparison, Mistral Large 3 lists at $0.50 input, $0.05 cached and $1.50 output, so Large 4 at list is about 2.7 times its predecessor on input and 2.8 times on output; at the sale price the gap is about 1.4 times. Mistral docs, pricing
Two things in the docs qualify the word “preview”. Mistral’s lifecycle policy says public preview models are “priced at the same rate as General Availability models”, can receive updates before GA, have “no guaranteed path to General Availability” and may be retired before reaching it, with a one-month deprecation notice instead of the six months a GA model gets. Mistral docs, model lifecycle And the context window is not consistent across Mistral’s own surfaces: the docs say 1M, while the OpenRouter endpoint served by Mistral lists 524,288 tokens with output capped at 262,144, and Artificial Analysis records 524k. OpenRouter Artificial Analysis Test the window you plan to use rather than the one on the model page.
1T or 1.05T
The launch post says 1 trillion total parameters. The docs page, versioned v26.10, says “52B active parameters and 1.05T total parameters, and a 1.6B vision encoder”. Mistral Mistral docs, Mistral Large 4 The post is probably rounding. Fifty billion parameters is a rounding error at this scale and 100 GB of BF16 weights when you are sizing a node, so the estimates below use 1.05T and show where 1T would change them.
Mistral has not published the layer count, expert count, experts per token or attention design. The post says those details, along with “additional benchmarks, and our post-training methodology”, will come with the weights. Mistral Our mixture-of-experts explainer covers why total and active counts answer different questions.
The license is a blank badge, and Mistral’s record cuts both ways
Nothing Mistral has published for Large 4 names a license. The launch post uses “open weights” and “open-weight” repeatedly and never says under what terms. The docs page’s badge reads “Open”, which on the same models list sits next to “Apache 2.0” for Large 3 and Small 4, “Modified MIT” for Medium 3.5 and “Open” for Z.ai’s GLM 5.3, a third-party model with a custom license of its own. Mistral docs, models Artificial Analysis, which tracks weights availability, lists Large 4 Preview as proprietary, with weights not publicly available. Artificial Analysis On Hugging Face, the mistralai organization’s model list shows no Large 4 repository. Hugging Face, mistralai
The precedent is mixed. Mistral Large 3, released December 2, 2025, is Apache 2.0: the announcement says “All models are released under the Apache 2.0 license”, and the Hugging Face repository for Mistral-Large-3-675B-Instruct-2512 carries that tag. Mistral, Mistral 3 Hugging Face Mistral Medium 3.5, docs version v26.04, is not. Its LICENSE file is a “Modified MIT License” whose second clause says you are not authorized to use the model, or any derivative, if your company’s “global consolidated monthly revenue” exceeded $20 million in the preceding month, unless Mistral grants a commercial license or you use its hosted service. Hugging Face, Mistral Medium 3.5 LICENSE So within one year Mistral has shipped its flagship as plain Apache 2.0 and its next-largest model with a revenue gate. Large 4 could go either way, and the post gives no signal.
Recent open-weight launches show how much the term can cover. Qwen3.8-Max’s weights arrived under a custom license with a $50 million commercial threshold, and the checkpoint was a text-only core rather than the hosted product. GLM 5.3 split into an MIT Flash model and a flagship whose license sends model-as-a-service operators above $10 billion through a security review. Reflection’s Beam has promised Apache 2.0 with no text to read. Until a LICENSE file exists in a Large 4 repository, “open-weight” is a description of intent.
Mistral’s benchmarks, and the one independent number
Mistral’s headline coding figures are DeepSWE v1.1 61.7%, SWE-Atlas-QnA 59.4% and Terminal-Bench 4 at 28.3%, with a Coding Agent Index of 49.8% that it says places the model ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. A blind human evaluation run with Surge AI put “ML4 Preview” second of five at 3.74 on a five-point scale, behind Claude Opus 5 (4.22) and ahead of GLM-5.3 (3.60), Kimi K3 (3.59) and GLM-5.2 (3.40). On agents, Mistral cites 59.9% on AutomationBench and 1,393 Elo on AA-Briefcase; on vision, 42% on Dense 200 against 41% for GPT-6 Astra. An internal human evaluation against GLM-5.3 preferred Large 4 “in CAD and STEM, while performing on par or close to GLM-5.3 in finance and coding.” Mistral The post’s own claim is “competitive with the strongest open-source models globally” and “significantly outperforming any open-weight model developed in the US or Europe”. Its human eval has the model behind Opus 5, and the comparison set for the open claim is Chinese models it describes as roughly level.
Artificial Analysis had evaluated the preview by October 7. Its page gives Large 4 Preview an Intelligence Index of 38 (v4.3.2), above the median of 26 for comparable models, at 116 output tokens per second. Two of its figures matter for cost. The model generated 200 million tokens across the index run, “very verbose in comparison to the median of 81M”, and the resulting cost per index task is $1.13 at list price. Artificial Analysis A model that is cheap per token but verbose per task is the case our Sonnet 5.5 piece made for measuring cost per finished task. Mistral’s cyber claims rest on the same evaluator: it says the model “ranks among the top five models globally” on the Artificial Analysis Cyber Index and scores 82% on the index’s reproduce-and-patch test. We could not read the model’s cyber rank from the index page, whose text lists only the top three (Grok 4.7, MiMo-V2.6-Pro and GPT-6 Luna). Artificial Analysis, Cyber Index Treat the top-five claim as Mistral’s until the page confirms it.
Cyber access is the same model with fewer refusals
The post is direct about why it leads with cybersecurity. It says Claude Opus 5.5 and GPT-6 Astra “score near zero” on the reproduce-and-patch test “because they refuse to perform the task”, and that defenders “need systems that can match those capabilities without being constrained by the same refusals.” Between now and the weights release, Mistral says it is “red-teaming the model in real-world settings with cybersecurity leaders, vetted partners, and state authorities, who will access the same model with reduced moderation and expanded cyber capabilities.” Mistral That is one model with two moderation settings, not a separate cyber build, and the post says the open weights will let organizations run security work “under their own policies”.
The same post reports a cyber refusal rate “higher than all OSS models” on JailbreakBench, StrongREJECT and AgentHarm prompts, so the public preview is tuned to refuse more than the partner build. Which behavior the released checkpoint carries is not stated. Anthropic’s answer to the same tension, published the same day, is a verification program that turns down cyber classifiers by tier for vetted applicants on hosted models; our coverage of that program is the useful comparison, because Mistral’s version of tiered access ends when the weights are public.
What a 1T, 52B-active checkpoint needs to run
No checkpoint exists, so these are estimates. Assumptions: decimal gigabytes, weights only, 1.05T parameters including the vision encoder, no KV cache, activations or runtime buffers. Precision figures use Mistral’s own Large 3 checkpoints as the ratio: the BF16 repository is 1,352 GB for 675B parameters (2.0 bytes per parameter), the default FP8 release 681.5 GB (1.01 bytes) and the NVFP4 build 403.1 GB, about 4.8 bits per parameter once scales are counted. Hugging Face, BF16 Hugging Face, FP8 Hugging Face, NVFP4
| Precision (estimate) | Weights at 1.05T | Weights at 1T | Published node or box that holds it |
|---|---|---|---|
| BF16, 2 bytes/param | ~2,100 GB | ~2,000 GB | Nothing in one node: DGX B300 lists 2.1 TB total GPU memory, which is the weights with no cache NVIDIA |
| FP8, ~1 byte/param | ~1,060 GB | ~1,010 GB | 8× H200 (1,128 GB) with roughly 70 GB spare NVIDIA |
| NVFP4, ~4.8 bits/param | ~630 GB | ~600 GB | 8× H200 comfortably; 8× H100 (640 GB) holds the weights with almost nothing left |
| Ideal 4-bit, 0.5 byte/param | ~525 GB | ~500 GB | A 512 GB Mac Studio M5 Ultra does not hold 1.05T at 4 bits; it would at 1T with no cache |
The last row matters most to anyone pricing a desktop. Mistral said Large 3’s NVFP4 build runs “on a single 8×A100 or 8×H100 node”. Mistral, Mistral 3 Scaled by parameter count, a Large 4 NVFP4 build does not leave room for a cache on 8× H100 and needs the H200 generation or better. A 512 GB unified-memory desktop, which Apple has said ships in late October, is below the ideal 4-bit figure at 1.05T; running the model there means a quantization below 4 bits, with the quality loss that implies, and it would be a single-user box in any case. Four 128 GB DGX Spark units reach 512 GB on paper with the same shortfall plus an interconnect penalty. The Open Models guide separates file size from resident memory and the machine guide covers headroom by box.
Per generated token, the model reads its active parameters, about 104 GB at BF16 or roughly 31 GB at NVFP4, which is a bandwidth figure, not a capacity one, and is why a 52B-active model can be fast once it fits.
The KV cache is the unknown. Mistral has not published the attention design, and the cache size depends on it entirely. Large 3 uses multi-head latent attention with a 512-dimension latent plus a 64-dimension rotary component across 61 layers, which works out to about 70 KB per token at BF16, or around 18 GB for a 256K context and about 74 GB at 1M. Hugging Face, Large 3 params If Large 4 keeps that family of attention, full-context cache is tens of gigabytes per sequence; if it uses conventional grouped-query attention at this width, it could be several times more. Our KV cache explainer covers the arithmetic. Until params.json is public, size the node for the weights at your precision plus at least the Large 3 cache figure at your context, then add concurrency.
What to check when the weights land
- A
mistralai/Mistral-Large-4-*repository on Hugging Face, whether it is gated, and the LICENSE file’s text rather than the badge. Apache 2.0 puts it with Large 3; a modified MIT with a revenue clause puts it with Medium 3.5. Check that the metadata tag and the file agree. - Which checkpoint ships: instruct, base, or both; whether the released model carries the public preview’s moderation tuning or the partner build’s; whether the vision encoder is included. Qwen3.8-Max’s download was not the hosted product, and the launch post calls Large 4 the base for “a new generation of specialized and optimized Mistral models”, so read the card for what was held back.
params.json: total parameters against 1.05T, experts and experts per token, layers, attention type andmax_position_embeddings(1M or 524,288).- Whether BF16, FP8 and NVFP4 builds all appear at launch, as they did for Large 3, and whether the sizes land near the estimates above.
- The architecture and post-training notes the post promises, with eval settings for the headline numbers: effort, harness, pass@k, run counts, and the five-model set behind the Surge AI human eval.
- Independent entries: Artificial Analysis moving the model from proprietary to open weights and publishing its cyber rank; third-party hosts on OpenRouter beyond Mistral itself; a vLLM or SGLang recipe with a real memory footprint at a stated context.
- Whether the half-price preview rate survives GA, and the one-month notice window if the preview is retired or replaced.
Until the repository exists, Large 4 is a priced API with a strong vendor benchmark page and one independent index score. It is worth a preview-API evaluation now if your workload is coding, agents or document vision, with cost measured per finished task given the verbosity Artificial Analysis recorded. It is not yet something to plan hardware or a license review around, and a 1.05T checkpoint will not fit the hardware most teams have on hand at any precision Mistral has shipped before.
Key Details
| Spec | Detail |
|---|---|
| Lab | Mistral AI (Paris) |
| Announced | October 6, 2026; public preview API on Mistral Studio |
| Weights | Promised “end of this month”; no Hugging Face repository as of October 7 |
| License | Not named; docs badge “Open”. Large 3 was Apache 2.0; Medium 3.5 was modified MIT with a $20M monthly-revenue gate |
| Total parameters | 1T (launch post) or 1.05T (docs, v26.10), self-reported |
| Active parameters | 52B per token (about 5%) |
| Vision encoder | 1.6B (docs only) |
| Context | 1M (docs); 524,288 with 262,144 max output on the OpenRouter endpoint served by Mistral |
| Model ID | mistral-large-4 (docs); mistral-large-4-0 on OpenRouter |
| Price (per 1M tokens) | List $1.36 in, $0.14 cached, $4.18 out; sale $0.68 / $0.07 / $2.09, no end date |
| Preview terms | Same price as GA, updates possible, one-month deprecation notice, no guaranteed GA |
| Training | 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s European datacenters; 160+ languages (self-reported) |
| Vendor benchmarks | DeepSWE v1.1 61.7%; Terminal-Bench 4 28.3%; Cybench 93% of 40; Dense 200 42%; Surge AI human eval 3.74, second of five |
| Independent | Artificial Analysis Intelligence Index 38 (v4.3.2), 116 tokens/s, $1.13 per index task, 200M output tokens in the run |
| Self-host (estimates) | ~2,100 GB BF16; ~1,060 GB FP8; ~630 GB NVFP4; KV cache unknown until the architecture is public |
Sources
- Introducing Mistral Large 4 — Mistral AI
- Mistral Large 4 model page — Mistral docs
- Models overview — Mistral docs
- Pricing — Mistral docs
- Model lifecycle policy — Mistral docs
- Introducing Mistral 3 — Mistral AI
- Mistral-Large-3-675B-Instruct-2512, BF16 and NVFP4 — Hugging Face
- Mistral-Medium-3.5-128B LICENSE — Hugging Face
- mistralai organization — Hugging Face
- Mistral Large 4 Preview — Artificial Analysis
- Artificial Analysis Cyber Index — Artificial Analysis
- mistralai/mistral-large-4-0 — OpenRouter
- H200 Tensor Core GPU and DGX B300 — NVIDIA
