Mistral Large 4 is the biggest model Mistral has ever built, and the French lab is betting that its open weights will matter most to organizations that cannot afford to depend on someone else’s safety filters. The model is available now as a public preview through Mistral Studio and the API, with full weights scheduled for release at the end of October. Internally the team calls it ML4. Their unofficial nickname is “le Chonk.”

This article covers what Mistral Large 4 is, where it performs well, where it falls short and why its approach to cybersecurity stands out among frontier models.

Mistral Large 4 specs at a glance

ML4 is a sparse mixture of experts model that combines instruction following, reasoning and agentic behavior in a single system. You no longer need to switch between a chat model and a separate reasoning model for different tasks.

  • Parameters 1 trillion in total, with 49 billion active per token
  • Context window 512k tokens, with up to 256k output tokens
  • Modalities text and image input, text output
  • Images per request up to 100, compared with 8 on earlier Mistral models
  • Languages trained on more than 160, including every official EU language
  • Tooling tool calling and structured outputs supported
  • Pricing $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 and a 50% launch discount during the first two weeks

For comparison, Mistral Large 3 used 675 billion total parameters and 41 billion active ones. The total parameter count grew considerably, while the active count rose only modestly, which keeps inference costs per token in a reasonable range.

Cybersecurity is the headline use case

Mistral puts cybersecurity front and center, and the benchmark results support that choice. On the Artificial Analysis Cyber Index, an independent evaluation of how well models find and fix flaws in real software, ML4 scores 50. That places it level with GLM-5.3-Flash and behind MiMo-V2.6-Pro at 56. Once the weights are public, Artificial Analysis expects it to rank among the top three open weights models on that index.

Its strongest result comes from CyberGym-E2E-AA, a test where the model has to reproduce a real vulnerability in open source software and then patch it. ML4 scores 82% there, the highest of any model tested. It also solves 93% of the 40 Cybench challenges drawn from security competitions.

Why refusals are a security problem

The CyberGym result reveals something important. Several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score close to zero on that same test because they refuse the task outright. Proving that a vulnerability is real is often the first step in defending against it, so a model that declines to reproduce exploits leaves defenders stuck.

Mistral argues this gap becomes riskier as attackers jailbreak those same closed models for offensive work. Defenders end up with weaker tools than the people they are fighting. Losing API access in the middle of an incident response is a second risk. With open weights and self deployment, a security team can run ML4 on a private cloud or on premises and apply its own policies.

In internal testing, ML4 also proved useful for malware analysis, vulnerability prioritization and writing detection rules, which are tasks it was never explicitly trained on. Before the public weight release, Mistral is red teaming the model with cybersecurity leaders, vetted partners and state authorities, who get access to a version with reduced moderation and expanded cyber capabilities.

Permissiveness for defenders does not mean the model helps everyone. Its average refusal rate on malicious cyber prompts from JailbreakBench, StrongREJECT and AgentHarm is higher than that of any other open source model. On Lakera’s B3 AI Security Benchmark, it resists 93.3% of indirect prompt injection attacks, the best score among competitors Mistral tested.

Coding and agentic workflows

ML4 scores 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4. Its combined Coding Agent Index score of 49.8% puts it ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max.

Mistral also ran a blind evaluation with Surge AI, where professional annotators rated coding outputs on a scale of 1 to 5 without knowing which model produced them. ML4 came second with 3.74, ahead of GLM-5.3, Kimi K3 and GLM-5.2. Only Claude Opus 5 scored higher, at 4.22.

For general business agents, AutomationBench is a useful reference. It contains 657 workflows across apps such as Gmail, Google Sheets, Slack and Salesforce. ML4 scores 59.9%, beating Kimi K3, MiMo-V2.6-Pro and DeepSeek V4 Pro. On AA-Briefcase, which measures long horizon knowledge work involving spreadsheets, slides and PDFs, it reaches 1,393 Elo.

Vision that can zoom, inspect and verify

Image understanding is where ML4 improves most over its predecessor. On GDP.pdf, a document and image reasoning test, it scores 19%, an 18 point jump from Mistral Large 3. Part of that gain comes from the larger image limit per request.

Visual grounding is the standout. On Dense 200, ML4 edges out GPT-6 Astra with 42% against 41%. Mistral pairs this with agentic behavior. The model can scan gigapixel satellite imagery for small objects during disaster response, or inspect an engineering drawing by zooming into specific parts until it can verify an answer. Manufacturing, engineering and earth observation teams are the obvious audience here.

Science, law and finance

On SciCode-Verified, which tests whether a model can implement scientific workflows in code across physics, mathematics, materials science and biology, ML4 is the best open weights model. Mistral reports that it can generate a full Hartree–Fock simulation in a single attempt, a multistep chemistry task built from a chain of advanced routines. In expert comparisons against GLM-5.3, annotators preferred ML4 for CAD and STEM work, while results in finance and coding were close.

For professional services, third party evaluator vals.ai found ML4 outperforms GPT-6 Astra on representative legal and financial tasks. On Harvey’s Legal Agent benchmark, it beats every open source model. FinWorkBench, which tests the creation and editing of spreadsheets for real accounting cases, shows similar strength.

The cost question

Price is ML4’s clear weakness. According to Artificial Analysis, it scores 38 on their Intelligence Index, comparable to GPT-6 Luna at 38 and DeepSeek V4.1 Flash at 39. That makes it the most intelligent model built outside the US and China. Running the full index costs $1.13 per task, though, which is over four times as much as open weights models with similar intelligence. GLM-5.3-Flash costs $0.25 per task and DeepSeek V4.1 Flash costs $0.27.

The launch discount brings ML4 down to $0.57 per task, which is still more than double. If you are running high volume general workloads where cyber performance, sovereignty or visual grounding are irrelevant, cheaper alternatives deliver similar raw intelligence. The case for ML4 rests on its specialist strengths and on where and how it can be deployed.

Built in Europe, trained with reinforcement learning at scale

Mistral trained ML4 from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters, and the preview runs on that same hardware. The model will be offered in several regions, including a European deployment that Mistral operates end to end under European law and independently of other digital service providers. For public sector bodies and regulated industries, that independence often matters more than a few benchmark points.

Post training relies heavily on reinforcement learning. Mistral’s RL library lets one training run combine many environments, from simple chat to scientific problem solving, safety alignment and long tool use sequences. An autoscaling fleet generates tens of thousands of rollouts in parallel, producing about 33 billion tokens per day, of which roughly 16 billion are usable for training. Customers get access to that same training and customization environment through Mistral Forge.

The RL run is still in progress, and Mistral says the model shows no signs of saturation. ML4 is also the first milestone funded by the company’s €3 billion Series D, the largest equity round ever raised by a European tech company, and much of that money is going into additional compute.

What to watch before the weights arrive

The preview you can test today is a snapshot of a model that is still training. Benchmark numbers will likely shift by the time the weights land at the end of October, and Mistral has promised details on the architecture and post training methods alongside them. The more interesting signal will come afterward, when independent security teams start running ML4 on their own infrastructure. Real incident response work will show whether the balance between helping defenders and refusing attackers holds up outside a benchmark suite. ML4 will also serve as the base for a new line of specialized Mistral models, so its strengths and its pricing will shape the rest of the lineup.