Lotu Radar About

Mistral’s new AI tried to escape its test environment. In three weeks, anyone can download it

The New Stack Cloud & Infrastructure Score 7/10
Mistral’s new AI tried to escape its test environment. In three weeks, anyone can download it

Summary

Mistral launched Large 4 on Tuesday, its first major model release since Medium 3.5 at the end of April and The post Mistral’s new AI tried to escape its test environment. In three weeks, anyone can download it appeared first on The New Stack .

Original Text

Mistral launched Large 4 on Tuesday, its first major model release since Medium 3.5 at the end of April and the company’s latest attempt to close the gap with China’s leading open-weight models.

The company says the new model is particularly strong in cybersecurity, and those capabilities surfaced during evaluation when the model tried to go beyond its testing environment, Mistral VP of Science Pierre Stock told Reuters, adding that the behavior was expected and that the company contained it using software.

The episode has not slowed Mistral’s release plans, and Large 4, nicknamed “Le Chonk” in a nod to the “Le Chaton Fat” meme about a fictional supersized Mistral model that spread across X and Reddit in June, is now in public preview through Mistral’s API, with cybersecurity experts and government authorities testing a version with fewer safety restrictions before the weights are published on October 27.

OpenAI and Anthropic have seen similar behavior while testing their most cyber-capable models and responded by restricting access. Mistral is taking a different route, with plans to release the Large 4 checkpoint in three weeks under a custom license rather than the Apache 2.0 license used for Large 3. Once the weights are out, developers control how the model runs and what safeguards they put around it.

Mistral is taking a different route, with plans to release the Large 4 checkpoint in three weeks under a custom license rather than the Apache 2.0 license used for Large 3.

One trillion, 49 billion active

Large 4 gets its one-trillion-parameter size from a sparse mixture-of-experts architecture that activates 49 billion parameters during inference, a significant step up from the 675 billion total and 41 billion active parameters in Large 3.

The company trained Large 4 from scratch in roughly two months on about 4,000 Nvidia Grace Blackwell GPUs in its European data centers. Mistral has made a point of the relatively small training cluster, although comparisons with the largest U.S. labs are difficult when so little training compute is disclosed.

Its sparse architecture keeps inference compute down by activating only 49 billion of the model’s one trillion parameters, but serving the full checkpoint will still require a substantial multi-GPU setup.

Software engineering and cybersecurity are the main targets for Large 4, with financial analysis, satellite and aerial imagery, technical drawings and chip design among its other use cases. It takes multimodal inputs, produces text and supports more than 160 languages, including every official language of the European Union.

Open weights, no recall

Mistral’s case for open weights in cybersecurity rests on control. Security teams scanning code or testing systems can run into a hosted model’s safety restrictions, and OpenAI’s safety system is already cutting off API responses mid-task even as the company gives models more authority inside its own development workflow, including blocking code from merging when a vulnerability is found.

Mistral’s case for open weights in cybersecurity rests on control.

Running Large 4 on their own infrastructure lets teams set those restrictions themselves and keep sensitive code and data in-house. Stock made the other half of that argument to Journal du Net, noting that once weights are replicated across the internet, access can no longer be easily revoked.

DeepSWE scores need context

Mistral reports a 62% score on DeepSWE v1.1, just above the 61% it lists for GLM-5.3. However, the live DeepSWE leaderboard puts GLM-5.3 and Kimi K3 at roughly 69% with their best published configurations, while GPT-6 Astra, Gemini 3.8 Flash and Claude Opus 5 are around 74%.

The results are stronger elsewhere. Large 4 reached a 15% task-pass rate on Harvey’s Legal Agent Benchmark and 67% on Finch, where Mistral’s testing has it tied with DeepSeek V4 Pro 0813 and ahead of GLM-5.3 at 65%. Independent results for those configurations aren’t yet available.

For now, those numbers make Large 4 look competitive without putting it at the top of the pack. We’ll all be eager to see what comes at the end of the month when developers can run Le Chonk outside Mistral’s API and see how much of that performance survives on their own workloads.

Large 4 reached a 15% task-pass rate on Harvey’s Legal Agent Benchmark and 67% on Finch, where Mistral’s testing has it tied with DeepSeek V4 Pro 0813 and ahead of GLM-5.3 at 65%.

The post Mistral’s new AI tried to escape its test environment. In three weeks, anyone can download it appeared first on The New Stack.

CloudInfrastructure

Lotu Radar provides attributed news summaries and links to the original publisher. Full reporting and copyright remain with the source.