Mistral Introduces 'Large 4' as a European Frontier Model, Bringing Advanced Vision and Open Deployment Closer to Reality

A model that can read a drawing, patch a flaw, and finish a spreadsheet is not new. What changed this week is who can run that loop without asking a provider for permission first.

European labs have spent the past few years arguing that frontier models can be trained and served without leaving the continent. The argument usually arrives as a policy paper.

Now, Mistral has put a model behind it.

'Mistral Large 4' entered public preview under the nickname "Le Chonk."

Mistral describes it as a granular mixture-of-experts model with roughly 1 trillion parameters, of which 49 billion are active, according to the launch post. Its documentation lists 1.05 trillion total parameters, 52 billion active, plus a 1.6 billion parameter vision encoder.

The model accepts text and images and returns text. It combines instruction following, reasoning, and tool use in a single checkpoint rather than separating them into different modes.

The weights are scheduled to arrive at the end of October. Until then, the open-weight claim comes with a date, while the full architecture writeup is still to come.

The previous flagship is the clearer before-and-after. 

Mistral Large 3, released in December 2025, was already a granular mixture-of-experts with a vision encoder built in, not bolted on later. It had 675 billion parameters, 41 billion active, a 2.5 billion parameter vision encoder, a 256,000 token context, and a limit of eight images per prompt. 

The weights shipped under Apache 2.0. 

Large 4 is not a fine-tune of that model. 

Instead, Mistral says it was trained from scratch, on 3,800 Nvidia's Grace Blackwell GPUs in the company's own European datacenters, over a run that press briefings put at about two months. 

Active parameters move only a little, from 41 billion to about 49 or 52 billion. Total parameters jump by roughly half. The vision encoder is smaller on the docs card, 1.6 billion against 2.5 billion, which suggests the gain is not a bigger eye but a language model that can use what the eye sees. 

Mistral calls image understanding a step change against its own earlier models. 

Artificial Analysis records one concrete version of that: the API now accepts 100 images in a request, up from eight, and on the GDP.pdf document task Large 4 scores 19%, an 18 point gain on Large 3.

The method Mistral is willing to describe is mostly post-training. 

Pretraining details are thin. What the company does spell out is a reinforcement-learning setup built because a recipe tuned for the last model stops pushing the next one. 

A shared interface lets one training run mix single-turn chat, scientific problem solving, safety, factuality, and long tool-use trajectories. Those environments share sandboxes, web search, and external APIs. 

Checks are mixed per task: reward models, unit tests, model judges, static tests. At run time an autoscaling pool produces tens of thousands of attempts in parallel while training continues asynchronously, with rollout budgets in the millions of tokens and extra steps to keep the policy from drifting. 

The same environment is what Mistral sells to customers as Forge, and the company says it trained alongside firms in finance, engineering, manufacturing, logistics, pharmaceuticals, and the public sector. 

A large share of the data is multilingual, more than 160 languages, including every official language of the European Union. The model is still being refined. Mistral says the pace from here will be fast, which is another way of saying the preview is not the final checkpoint.

The capability list follows from that mix. Against the established American systems, the gap is not uniform, and the interesting part is where the score is a policy score. 

On the blind coding test, Claude Opus 5 still leads, at 4.22 against 3.74. On visual grounding the edge over GPT-6 Astra is one point, 42 against 41. On third-party legal and finance tasks from Vals, Mistral says Large 4 exceeds GPT-6 Astra, and on Harvey's legal agent benchmark it leads the open models. 

The sharp difference is cybersecurity. 

On a test that asks a model to reproduce a real flaw in open-source software and then patch it, Large 4 scores 82%, which Mistral calls the highest of any model, and 93% of Cybench's 40 competition exercises. 

Claude Opus 5.5 and GPT-6 Astra score near zero on that reproduce-and-patch task because they refuse it. 

Defensive work often starts by proving a bug is real. A provider filter can block that in the middle of an incident, and Mistral's argument is that a weight you host yourself can be moderated under your own policy instead. 

The company is red-teaming the preview with security firms and state bodies on a build with reduced moderation. 

It also reports a higher refusal rate than other open models on jailbreak-style cyber prompts, so the public checkpoint is not an unfiltered tool. On Lakera's B3 injection benchmark it resists 93.3% of attacks, and on KORA, a behavior score with 2 as the top mark, it sits at 1.691.

Large 3 remains the model you can download today, with a known license and a smaller context. Large 4 is the one Mistral is pointing at drawings, filings, and defensive code, with the American comparison resting less on a general leaderboard than on tasks those labs currently decline to run.

In the end, the bigger shift is therefore not simply that Mistral has made a larger model. 

It is that the company is moving toward a model that can be hosted, inspected, and adapted on European infrastructure, with fewer restrictions imposed by a remote API. Large 4 is still a preview, and its weights have not arrived yet. 

But if the October release holds, "open" will become a much more practical claim than it was when the announcement first landed.

Published