AI & Machine Learning

Mistral Large 4: a trillion parameters from Europe, and what "open" does not mean yet

Mistral Large 4 is the first European open-weight model with a trillion parameters. What Mistral's figures do and do not say, what it costs, what hardware you need and why the licence is the biggest open question.

Erik van de Blaak
Erik van de Blaak
10 min read • 11 views
Mistral Large 4: a trillion parameters from Europe, and what "open" does not mean yet

If you want to run Mistral Large 4 on your own server, you need about 525 gigabytes of memory at 4-bit precision. For the weights alone. That is the memory of seventeen RTX 5090s. Even so, this is the model Europe was waiting for: the first European open-weight model with a trillion parameters, trained in European data centres and launched as a preview on 6 October 2026.

Mistral calls it ML4 internally, and "very officially: le Chonk". The announcement is full of charts in which the model wins. Read past the headlines and you find a more modest and more interesting story. Large 4 is not the best model in the world, and Mistral does not claim it is. It is the strongest model you will soon be able to download outside China. For organisations that do not want their data with an American or Chinese provider, that changes the choice.

Facts checked on 9 October 2026. Large 4 is a preview: the figures come from Mistral or from the media named and may change before the final version. I did not test the model myself for this article. Back-of-envelope calculations are marked as such.

What is Mistral Large 4?

Mistral Large 4 is a language model from the French company Mistral AI with about a trillion parameters, 52 billion of which take part per token. It reads text and images, writes text, can answer directly or reason step by step first, and has a context window of one million tokens. Since 6 October it has been available as a preview through the API in Mistral Studio. The weights follow "by the end of the month" according to Mistral, on 27 October according to Tweakers and VentureBeat.

Mistral Large 4Preview, 6 October 2026
Parameters1.05 trillion according to the model card; Mistral rounds to 1 trillion
Active per token52 billion (early reports said 49 billion)
Designmixture of experts, with a 1.6-billion-parameter image encoder
Input and outputtext and images in, text out
Context window1 million tokens
Languagesmore than 160, including all official EU languages
Training3,800 Nvidia Grace Blackwell GPUs in European data centres
API price$1.36 per million input tokens, $4.18 per million output tokens; half price during the launch
Weightsend of October; licence not yet published

That sources contradict each other (49 or 52 billion active, 3,800 or "about 4,000" GPUs) mostly tells you where the model stands: Mistral is still training it. The reinforcement learning run is ongoing and, according to Mistral, the model shows "no signs of saturation". The company will publish the architecture and training method together with the weights.

A trillion parameters, computing like a 52-billion model

Large 4 is a mixture-of-experts model. The network consists of many sub-networks, the experts, and for each token a router picks a handful of them. Per word, the model therefore computes about as much as an ordinary 52-billion-parameter model, while carrying the knowledge of a trillion parameters. That explains how a model of this size can be sold for $4.18 per million output tokens.

Many reports skip the consequence: computing fast is not the same as being small. All experts must sit in memory, including the ones that do nothing for your question, because the router only decides per token which ones it needs.

Bar chart: memory for the weights of Mistral Large 4 alone. BF16 (2 bytes per parameter) 2,100 GB, FP8 (1 byte) 1,050 GB, 4-bit (half a byte) 525 GB. For comparison, vertical lines at an RTX 5090 (32 GB), a workstation with 256 GB of memory and a server with eight H200s (1,128 GB). Only the FP8 and 4-bit versions fit in the eight H200s.
Back-of-envelope: parameters times bytes per parameter, without the memory the context needs on top.

The calculation: 1.05 trillion parameters times 2 bytes in BF16 is 2.1 terabytes. In FP8 (1 byte per parameter) it is 1.05 terabytes, in 4-bit 525 gigabytes. Memory for the context comes on top, and if you actually fill the million-token window, you need a lot more. An RTX 5090 has 32 gigabytes. A server with eight Nvidia H200s has 1,128, enough for the FP8 version with a small margin.

For comparison: Mistral Large 3, released in December 2025, had 675 billion parameters, 41 billion of them active. Large 4 is more than one and a half times as large and does only a quarter more computation per token.

"Open" for this model therefore means: open to organisations with a server rack and to hosting companies that offer the model. Not to your laptop. Under the Tweakers news item, a large share of the comments was about exactly that: the popular open models run on a machine with 128 or 256 gigabytes of memory, and this one does not.

The Large 4 licence is still unknown

Mistral Large 3 was released under Apache 2.0. That lets you use the model commercially, modify it and host it for others without asking. For Large 4, Mistral has not published a licence yet. VentureBeat reports that the weights will come under a custom Mistral licence; Tweakers writes that it is unclear which one.

The difference matters. Mistral itself calls Large 4 "open-weight", not open source. In a custom licence a company can set all kinds of terms, such as a revenue limit for free commercial use, a ban on offering the model as a service to third parties, or conditions on derived models. Nobody outside Mistral knows what this licence will say. If you are designing an architecture today in which Large 4 runs on your own servers, you are building on a document that does not exist yet.

How good is Mistral Large 4? Its own figures are more honest than the headlines

The headlines say: the strongest open model outside China. Mistral's own charts show what that means.

On DeepSWE v1.1, a test in which a model acts as an agent and solves real programming tasks, Large 4 scores 61.7 percent in Mistral's chart. That is just above GLM-5.3 (61) and well above DeepSeek V4 Pro (57) and Qwen 3.8 Max (51). VentureBeat put the public DeepSWE leaderboard next to it. With their best settings, Claude Opus 5, GPT-6 Astra and Gemini 3.8 Flash reach about 74 percent there, GLM-5.3 and Kimi K3 about 69. The setups differ, so the numbers cannot be compared one to one. The gap with the closed top is larger than Mistral's chart suggests, though.

Bar chart DeepSWE v1.1, share of programming tasks solved. Mistral's chart: Mistral Large 4 61.7, GLM-5.3 61, DeepSeek V4 Pro 0813 57, Qwen 3.8 Max 51, Reflection Beam 44. Public leaderboard with the best setup per model, approximately: Claude Opus 5 74, GPT-6 Astra 74, Gemini 3.8 Flash 74, GLM-5.3 69, Kimi K3 69.
Top: the figures from the announcement. Bottom: the public leaderboard as VentureBeat looked it up. Large 4 was not on it yet at launch.

In a blind test by Surge AI, reviewers gave code from five models a score from 1 to 5. Large 4 got 3.74: second of the five, behind Claude Opus 5 with 4.22. Mistral publishes that figure itself, which speaks for the company. It does not hide that Anthropic's model writes better code.

Large 4 is stronger outside programming. On Harvey's legal agent benchmark it reaches 15 percent, against 13 for Kimi K3 and 5 for GPT-6 Astra, although several closed models rank higher on the Vals leaderboard. On DIOR-RSVG, where a model has to point out objects in satellite and aerial images, it scores 73 percent against 68 for GPT-6 Astra. That Mistral highlights exactly this test fits the customers it has in mind: VentureBeat lists aerial imagery, technical drawings and chip design among the target areas.

Cybersecurity: strong on a test others refuse

The most striking figure in the announcement is about security. On a test in which a model has to reproduce a vulnerability and then patch it, Large 4 scores 82 percent according to Mistral. Claude Opus 5.5 and GPT-6 Astra score "near zero" on the same test, Mistral writes, because they refuse the task.

That zero measures what the makers let their model do, not what it can do. Mistral turns the difference into a selling point. The company is testing Large 4 for three weeks with security firms, vetted partners and government agencies, who according to the announcement get "the same model with reduced moderation and expanded cyber capabilities". According to co-founder Guillaume Lample, quoted by TNW, this should help companies and governments defend themselves against attackers.

At the same time, Mistral claims that the public version refuses more often on three jailbreak tests (JailbreakBench, StrongREJECT and AgentHarm) than any other open model tested. Those claims need not clash: reproducing and patching a known flaw is different from carrying out an attack. They do show where the choice moves. With the big American models, the provider draws the same line for everyone. With Large 4, Mistral decides who gets a wider line. And once the weights are public, anyone who downloads them can move that line with their own training.

What Mistral Large 4 costs

Through the API, Large 4 costs $1.36 per million input tokens and $4.18 per million output tokens. During the launch Mistral charges half: $0.68 and $2.09. Input that is already in the cache costs $0.14 per million tokens, $0.07 during the offer.

A back-of-envelope example at the normal price: you put a codebase or case file of one million tokens in the context window and get a 50,000-token answer back. That costs $1.36 plus $0.21, $1.57 in total. Ask ten follow-up questions about the same material and the cache decides the bill. If the input comes from the cache, ten times a million tokens costs $1.40. Without the cache it is $13.60. If you work with long documents, first find out when Mistral's cache is and is not hit.

Who Large 4 does make a difference for

Until now, a European organisation that wanted to run a top model on its own servers mostly had Chinese options: DeepSeek, GLM, Kimi, Qwen. Most open-weight models come from China, Tweakers writes, and DeepSeek even has larger ones. Technically, the origin matters little: a model on your own server does not send data to Beijing. That argument came up again and again in the Tweakers comments, and it is correct.

What the argument misses is everything around the model: who trained it, on what data, under which law, and whom you call when something goes wrong. According to TNW, Mistral's customers want two guarantees: that no foreign government can cut off their access, and a clear record of how the model was trained. Mistral therefore offers Large 4 through the API in a chosen region, including a European environment the company says it runs "end-to-end" itself, in a private cloud, or on your own servers. According to TNW the company works for more than 125 large businesses, including Airbus, ASML and HSBC.

For a hospital, a municipality, a defence supplier or a bank, that is the new option: a model close to the top, with a European supplier and a copy of its own. For a developer who simply wants a good model through an API, Large 4 is one of many, and for programming not the strongest.

The money to hold that position is there. In September Mistral raised €3 billion at a valuation of more than €21 billion, according to the company the largest equity round ever for a European tech company. Large 4 is the first model from that round. According to TNW, the extra compute comes online in the first half of 2027.

What you can do with Mistral Large 4 now

Test it this month on your own work. The API costs half during the launch, and twenty real tasks from your own practice tell you more than a chart from the maker. Put the answers next to those of the model you use now, and write down where the difference is.

Only plan anything on your own servers once three things are known at the end of October: the licence text, the size of the checkpoint, and which smaller or quantised versions Mistral ships with it. Mistral also promises to publish the architecture and the training method then.

The question Large 4 answers is not whether Europe can build the best model in the world. It cannot yet, and Mistral's own charts show it. The question is whether a European organisation can own a model close to the top instead of renting one. At the end of October we will learn on what terms.

Sources

Share this article

Erik van de Blaak

Written by

Erik van de Blaak

AI Solutions Engineer & Full-Stack Developer

Erik van de Blaak is an AI Solutions Engineer and full-stack developer at CareerValue BV. Here he writes about AI coding tools and agents, such as Claude Code.

Comments (0)

Comment on this article

Never published

I read every comment before it goes online.

No comments on this article yet.

Read next