Mistral previews Large 4, a 1 trillion-parameter model, and says weights come by month's end
Mistral says Claude Opus 5.5 and GPT-6 Astra refuse a cyber test that ML4 scores 82% on.
On October 6, 2026, Mistral put out a public preview of Mistral Large 4, its largest model so far. The company calls it ML4 for short, and "very officially" le Chonk. You can call it through Mistral Studio today, and Mistral says it'll release the weights by the end of the month.
It's a 1 trillion-parameter model with 49 billion parameters active at a time, and it takes images as well as text. Mistral says it trained it from scratch in its own datacenters in Europe.
- Maker
- Mistral
- Size
- 1 trillion parameters, 49 billion active
- Trained on
- 3,800 NVIDIA Grace Blackwell GPUs
- Languages
- more than 160
- Preview price
- $1.36 input, $4.18 output per million tokens
- Weights
- by the end of the month
The cyber pitch
Most of the post is about security work. On one test in the Artificial Analysis Cyber Index, a model has to reproduce a real vulnerability in open-source software and then patch it. Mistral says ML4 scores 82% there, the highest of any model.

Closed models do far worse on that one test, and Mistral says why: they won't do it.
Several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same test because they refuse to perform the task.
Mistral's case is that defenders often have to prove a flaw is real before they can fix it, and that a provider's refusal in the middle of an incident is a risk in itself. With open weights, a security team can run the model on its own hardware under its own rules.
The same post also says ML4 refuses malicious cyber prompts more often than any other open model, averaged across JailbreakBench, StrongREJECT and AgentHarm. Until the weights are out, Mistral is red-teaming it with security firms, vetted partners and state authorities, and those testers get a version with "reduced moderation and expanded cyber capabilities."
How it compares
Mistral says ML4 beats any open-weight model built in the US or Europe, and is competitive with the strongest open models anywhere. Those are its own numbers. On coding it reports 61.7% on DeepSWE v1.1 and 28.3% on Terminal-Bench 4.
In a blind coding test run with Surge AI, professional annotators scored outputs from 1 to 5. ML4 Preview came second of five at 3.74, ahead of Kimi K3 and GLM-5.3, and behind Claude Opus 5 at 4.22. On one image benchmark, Dense 200, Mistral puts it a point ahead of GPT-6 Astra (42% to 41%).
The reinforcement learning run behind this preview is still in flight, and the model is showing no signs of saturation
What's still to come
The preview isn't the finished model. Mistral says the training run is still going, and it expects "large and rapid improvements" over the coming weeks. I'd hold off comparing it closely until the weights land and outside testers can run it.
Mistral calls ML4 the first milestone paid for by its €3 billion Series D, and says it'll be the base for a new set of specialized Mistral models. The architecture details, more benchmarks and the post-training method are due with the weights, by the end of October.
More on Open weights
- Reflection's first open-weight model, Beam, has 501 billion parameters, with weights due this monthOctober 5, 2026
- Alibaba says Qwen ran a month of self-improvement unattended, and plans a 10-trillion-parameter modelSeptember 23, 2026
- Qwen models explained: Qwen3.8 Max, Flash and 27B, open weights vs API, with release datesSeptember 23, 2026
- Intrinsic open-sourced the robot control stack it uses in real manufacturing deployments, under Apache 2.0September 22, 2026