Mistral Large 4: a trillion parameters from Europe, and what "open" does not mean yet
Mistral Large 4 is the first European open-weight model with a trillion parameters. What Mistral's figures do and do not say, what it costs, what hardware you need and why the licence is the biggest open question.

If you want to run Mistral Large 4 on your own server, you need about 525 gigabytes of memory at 4-bit precision. For the weights alone. That is the memory of seventeen RTX 5090s. Even so, this is the model Europe was waiting for: the first European open-weight model with a trillion parameters, trained in European data centres and launched as a preview on 6 October 2026.
Mistral calls it ML4 internally, and "very officially: le Chonk". The announcement is full of charts in which the model wins. Read past the headlines and you find a more modest and more interesting story. Large 4 is not the best model in the world, and Mistral does not claim it is. It is the strongest model you will soon be able to download outside China. For organisations that do not want their data with an American or Chinese provider, that changes the choice.
Facts checked on 9 October 2026. Large 4 is a preview: the figures come from Mistral or from the media named and may change before the final version. I did not test the model myself for this article. Back-of-envelope calculations are marked as such.
What is Mistral Large 4?
Mistral Large 4 is a language model from the French company Mistral AI with about a trillion parameters, 52 billion of which take part per token. It reads text and images, writes text, can answer directly or reason step by step first, and has a context window of one million tokens. Since 6 October it has been available as a preview through the API in Mistral Studio. The weights follow "by the end of the month" according to Mistral, on 27 October according to Tweakers and VentureBeat.
| Mistral Large 4 | Preview, 6 October 2026 |
|---|---|
| Parameters | 1.05 trillion according to the model card; Mistral rounds to 1 trillion |
| Active per token | 52 billion (early reports said 49 billion) |
| Design | mixture of experts, with a 1.6-billion-parameter image encoder |
| Input and output | text and images in, text out |
| Context window | 1 million tokens |
| Languages | more than 160, including all official EU languages |
| Training | 3,800 Nvidia Grace Blackwell GPUs in European data centres |
| API price | $1.36 per million input tokens, $4.18 per million output tokens; half price during the launch |
| Weights | end of October; licence not yet published |
That sources contradict each other (49 or 52 billion active, 3,800 or "about 4,000" GPUs) mostly tells you where the model stands: Mistral is still training it. The reinforcement learning run is ongoing and, according to Mistral, the model shows "no signs of saturation". The company will publish the architecture and training method together with the weights.
A trillion parameters, computing like a 52-billion model
Large 4 is a mixture-of-experts model. The network consists of many sub-networks, the experts, and for each token a router picks a handful of them. Per word, the model therefore computes about as much as an ordinary 52-billion-parameter model, while carrying the knowledge of a trillion parameters. That explains how a model of this size can be sold for $4.18 per million output tokens.
Many reports skip the consequence: computing fast is not the same as being small. All experts must sit in memory, including the ones that do nothing for your question, because the router only decides per token which ones it needs.
The calculation: 1.05 trillion parameters times 2 bytes in BF16 is 2.1 terabytes. In FP8 (1 byte per parameter) it is 1.05 terabytes, in 4-bit 525 gigabytes. Memory for the context comes on top, and if you actually fill the million-token window, you need a lot more. An RTX 5090 has 32 gigabytes. A server with eight Nvidia H200s has 1,128, enough for the FP8 version with a small margin.
For comparison: Mistral Large 3, released in December 2025, had 675 billion parameters, 41 billion of them active. Large 4 is more than one and a half times as large and does only a quarter more computation per token.
"Open" for this model therefore means: open to organisations with a server rack and to hosting companies that offer the model. Not to your laptop. Under the Tweakers news item, a large share of the comments was about exactly that: the popular open models run on a machine with 128 or 256 gigabytes of memory, and this one does not.
The Large 4 licence is still unknown
Mistral Large 3 was released under Apache 2.0. That lets you use the model commercially, modify it and host it for others without asking. For Large 4, Mistral has not published a licence yet. VentureBeat reports that the weights will come under a custom Mistral licence; Tweakers writes that it is unclear which one.
The difference matters. Mistral itself calls Large 4 "open-weight", not open source. In a custom licence a company can set all kinds of terms, such as a revenue limit for free commercial use, a ban on offering the model as a service to third parties, or conditions on derived models. Nobody outside Mistral knows what this licence will say. If you are designing an architecture today in which Large 4 runs on your own servers, you are building on a document that does not exist yet.
How good is Mistral Large 4? Its own figures are more honest than the headlines
The headlines say: the strongest open model outside China. Mistral's own charts show what that means.
On DeepSWE v1.1, a test in which a model acts as an agent and solves real programming tasks, Large 4 scores 61.7 percent in Mistral's chart. That is just above GLM-5.3 (61) and well above DeepSeek V4 Pro (57) and Qwen 3.8 Max (51). VentureBeat put the public DeepSWE leaderboard next to it. With their best settings, Claude Opus 5, GPT-6 Astra and Gemini 3.8 Flash reach about 74 percent there, GLM-5.3 and Kimi K3 about 69. The setups differ, so the numbers cannot be compared one to one. The gap with the closed top is larger than Mistral's chart suggests, though.
In a blind test by Surge AI, reviewers gave code from five models a score from 1 to 5. Large 4 got 3.74: second of the five, behind Claude Opus 5 with 4.22. Mistral publishes that figure itself, which speaks for the company. It does not hide that Anthropic's model writes better code.
Large 4 is stronger outside programming. On Harvey's legal agent benchmark it reaches 15 percent, against 13 for Kimi K3 and 5 for GPT-6 Astra, although several closed models rank higher on the Vals leaderboard. On DIOR-RSVG, where a model has to point out objects in satellite and aerial images, it scores 73 percent against 68 for GPT-6 Astra. That Mistral highlights exactly this test fits the customers it has in mind: VentureBeat lists aerial imagery, technical drawings and chip design among the target areas.
Cybersecurity: strong on a test others refuse
The most striking figure in the announcement is about security. On a test in which a model has to reproduce a vulnerability and then patch it, Large 4 scores 82 percent according to Mistral. Claude Opus 5.5 and GPT-6 Astra score "near zero" on the same test, Mistral writes, because they refuse the task.
That zero measures what the makers let their model do, not what it can do. Mistral turns the difference into a selling point. The company is testing Large 4 for three weeks with security firms, vetted partners and government agencies, who according to the announcement get "the same model with reduced moderation and expanded cyber capabilities". According to co-founder Guillaume Lample, quoted by TNW, this should help companies and governments defend themselves against attackers.
At the same time, Mistral claims that the public version refuses more often on three jailbreak tests (JailbreakBench, StrongREJECT and AgentHarm) than any other open model tested. Those claims need not clash: reproducing and patching a known flaw is different from carrying out an attack. They do show where the choice moves. With the big American models, the provider draws the same line for everyone. With Large 4, Mistral decides who gets a wider line. And once the weights are public, anyone who downloads them can move that line with their own training.
What Mistral Large 4 costs
Through the API, Large 4 costs $1.36 per million input tokens and $4.18 per million output tokens. During the launch Mistral charges half: $0.68 and $2.09. Input that is already in the cache costs $0.14 per million tokens, $0.07 during the offer.
A back-of-envelope example at the normal price: you put a codebase or case file of one million tokens in the context window and get a 50,000-token answer back. That costs $1.36 plus $0.21, $1.57 in total. Ask ten follow-up questions about the same material and the cache decides the bill. If the input comes from the cache, ten times a million tokens costs $1.40. Without the cache it is $13.60. If you work with long documents, first find out when Mistral's cache is and is not hit.
Who Large 4 does make a difference for
Until now, a European organisation that wanted to run a top model on its own servers mostly had Chinese options: DeepSeek, GLM, Kimi, Qwen. Most open-weight models come from China, Tweakers writes, and DeepSeek even has larger ones. Technically, the origin matters little: a model on your own server does not send data to Beijing. That argument came up again and again in the Tweakers comments, and it is correct.
What the argument misses is everything around the model: who trained it, on what data, under which law, and whom you call when something goes wrong. According to TNW, Mistral's customers want two guarantees: that no foreign government can cut off their access, and a clear record of how the model was trained. Mistral therefore offers Large 4 through the API in a chosen region, including a European environment the company says it runs "end-to-end" itself, in a private cloud, or on your own servers. According to TNW the company works for more than 125 large businesses, including Airbus, ASML and HSBC.
For a hospital, a municipality, a defence supplier or a bank, that is the new option: a model close to the top, with a European supplier and a copy of its own. For a developer who simply wants a good model through an API, Large 4 is one of many, and for programming not the strongest.
The money to hold that position is there. In September Mistral raised €3 billion at a valuation of more than €21 billion, according to the company the largest equity round ever for a European tech company. Large 4 is the first model from that round. According to TNW, the extra compute comes online in the first half of 2027.
What you can do with Mistral Large 4 now
Test it this month on your own work. The API costs half during the launch, and twenty real tasks from your own practice tell you more than a chart from the maker. Put the answers next to those of the model you use now, and write down where the difference is.
Only plan anything on your own servers once three things are known at the end of October: the licence text, the size of the checkpoint, and which smaller or quantised versions Mistral ships with it. Mistral also promises to publish the architecture and the training method then.
The question Large 4 answers is not whether Europe can build the best model in the world. It cannot yet, and Mistral's own charts show it. The question is whether a European organisation can own a model close to the top instead of renting one. At the end of October we will learn on what terms.
Sources
- Mistral AI, Mistral Large 4 (announcement, 6 October 2026)
- Mistral AI, Mistral Large 4 model card: parameters, context window and prices (consulted 9 October 2026)
- Tweakers, Europese Mistral brengt openweightmodel Large 4 met biljoen parameters uit (6 October 2026, Dutch), with the comments below it
- VentureBeat, Mistral debuts Large 4 'Le Chonk': licence, comparison with Large 3 and with the public DeepSWE leaderboard
- The Next Web, Europe's Mistral launches Large 4 to challenge China's lead in open AI models: training, customers and quotes
- Orrick, Mistral's €3B Series D at €21B post-money valuation (September 2026)


