Mistral opens the preview, and leads with the capability rivals withhold

Mistral AI launched a public preview of Mistral Large 4 on 6 October, a model it calls ML4 and, less formally, “le Chonk”. It is a natively multimodal mixture-of-experts model with 1 trillion total parameters and 49 billion active for any single token, and Mistral says it is the company’s largest and most capable model to date. The weights are not out yet: Mistral said it will publish them at the end of October, once red-teaming finishes.

Until then access is split in two. The preview API is open on Mistral Studio at $1.36 per million input tokens and $4.18 per million output tokens. A second group — “cybersecurity leaders, vetted partners, and state authorities”, in Mistral’s own words — gets the same model with reduced moderation and expanded cyber capabilities.

Blue network cables plugged into a patch panel inside a server rack
The preview API is served on the same European infrastructure the model was trained on. Illustration. Brett Sayles · pexels · Pexels License

The pitch is a model that will not refuse

Mistral’s central claim is about security work, and it is unusually direct about who the comparison is with. On the Artificial Analysis Cyber Index, an independent evaluation of how well models find and fix flaws in real software, Mistral said ML4 ranks among the top five models globally. On the index test that asks a model to reproduce a real vulnerability in open-source software and then patch it, Mistral reported 82%, which it says is the highest score of any model.

The reason, by Mistral’s account, is not capability alone. The company said Claude Opus 5.5 and GPT-6 Astra “score near zero on the same test because they refuse to perform the task”. Its argument is that proving a flaw is real is where defence begins, and that a provider-level refusal in the middle of an incident is itself a security risk. On Cybench, a set of 40 exercises drawn from security competitions, Mistral reported 93% — a figure from its own testing, not an independent run.

The framing arrives days after Google shipped its first Gemini 4 model to cyber defenders with the cyber guardrails relaxed, and it points at the same unsettled question: who gets handed a model that will do offensive security work on request.

The rest of the numbers are mostly Mistral’s own

On coding, Mistral reported 61.7% on DeepSWE v1.1, 28.3% on Terminal-Bench 4 and 59.4% on SWE-Atlas-QnA, and a combined Coding Agent Index of 49.8% that it places ahead of DeepSeek V4 Pro and Qwen3.8 Max. On AutomationBench — 657 business workflows across tools including Gmail, Google Sheets, Slack and Salesforce — it reported 59.9%.

A blind human evaluation run with Surge AI placed ML4 second of five models on coding quality, at 3.74 out of 5, ahead of Moonshot AI’s Kimi K3 at 3.59 and GLM-5.3 at 3.60, and behind Claude Opus 5 at 4.22. On visual grounding Mistral said ML4 scores 42% on Dense 200 against 41% for GPT-6 Astra, the one place it claims to pass a frontier closed model.

Mistral’s own summary is carefully bounded: ML4 is “competitive with the strongest open-source models globally” while “significantly outperforming any open-weight model developed in the US or Europe”. That is a claim about a regional field rather than the global one, and the Chinese open-weight models remain the thing to beat.

Two colleagues standing in a bright office, looking through a document together
Mistral says a European deployment will run end to end under European law. Illustration. Yan Krukau · pexels · Pexels License

Trained in Europe, on 3,800 chips

Mistral said ML4 was trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in its own European data centres, and that the public preview is served on the same infrastructure. A significant share of the training data was multilingual, the company said, spanning more than 160 languages including every official language of the European Union, and a European deployment will run end to end under European law, independently of other digital service providers.

What is missing is everything that would let an outsider check the claims: the weights, the architecture, the post-training methodology and the licence terms. Mistral said all of it ships with the release at the end of the month. Until then every figure above is a vendor number, with the single exception of the Artificial Analysis index placing.