ThinkFacility Sign in
  1. Home
  2. News
  3. Open weights

Mistral previews Large 4, a 1 trillion-parameter model, and says weights come by month's end

Mistral says Claude Opus 5.5 and GPT-6 Astra refuse a cyber test that ML4 scores 82% on.

On October 6, 2026, Mistral put out a public preview of Mistral Large 4, its largest model so far. The company calls it ML4 for short, and "very officially" le Chonk. You can call it through Mistral Studio today, and Mistral says it'll release the weights by the end of the month.

It's a 1 trillion-parameter model with 49 billion parameters active at a time, and it takes images as well as text. Mistral says it trained it from scratch in its own datacenters in Europe.

Maker
Mistral
Size
1 trillion parameters, 49 billion active
Trained on
3,800 NVIDIA Grace Blackwell GPUs
Languages
more than 160
Preview price
$1.36 input, $4.18 output per million tokens
Weights
by the end of the month

The cyber pitch

Most of the post is about security work. On one test in the Artificial Analysis Cyber Index, a model has to reproduce a real vulnerability in open-source software and then patch it. Mistral says ML4 scores 82% there, the highest of any model.

A man in a dark suit, white shirt and dotted tie sitting on a stage with a small microphone at his cheek, purple lights behind him
Arthur Mensch, Mistral's co-founder and CEO, at Techarena 2026. Photo: Jan2342342423, CC0, via Wikimedia Commons

Closed models do far worse on that one test, and Mistral says why: they won't do it.

Several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same test because they refuse to perform the task.

From Introducing Mistral Large 4 | Mistral

Mistral's case is that defenders often have to prove a flaw is real before they can fix it, and that a provider's refusal in the middle of an incident is a risk in itself. With open weights, a security team can run the model on its own hardware under its own rules.

The same post also says ML4 refuses malicious cyber prompts more often than any other open model, averaged across JailbreakBench, StrongREJECT and AgentHarm. Until the weights are out, Mistral is red-teaming it with security firms, vetted partners and state authorities, and those testers get a version with "reduced moderation and expanded cyber capabilities."

How it compares

Mistral says ML4 beats any open-weight model built in the US or Europe, and is competitive with the strongest open models anywhere. Those are its own numbers. On coding it reports 61.7% on DeepSWE v1.1 and 28.3% on Terminal-Bench 4.

In a blind coding test run with Surge AI, professional annotators scored outputs from 1 to 5. ML4 Preview came second of five at 3.74, ahead of Kimi K3 and GLM-5.3, and behind Claude Opus 5 at 4.22. On one image benchmark, Dense 200, Mistral puts it a point ahead of GPT-6 Astra (42% to 41%).

The reinforcement learning run behind this preview is still in flight, and the model is showing no signs of saturation

Introducing Mistral Large 4 | Mistral

What's still to come

The preview isn't the finished model. Mistral says the training run is still going, and it expects "large and rapid improvements" over the coming weeks. I'd hold off comparing it closely until the weights land and outside testers can run it.

Mistral calls ML4 the first milestone paid for by its €3 billion Series D, and says it'll be the base for a new set of specialized Mistral models. The architecture details, more benchmarks and the post-training method are due with the weights, by the end of October.

More on Open weights

All Open weights stories