Power Your Software Testing with AI Agents and Cloud
The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.
- TestMu AI (Formerly LambdaTest)
- /
- Blog
- /
- Mistral Large 4 (Le Chonk): Specs, Benchmarks and Open Weights
Mistral Large 4 (Le Chonk): Specs, Benchmarks and Open Weights
Mistral Large 4 (Le Chonk) is a 1T-parameter open-weight multimodal model. Its specs, benchmarks, cyber results, open-weights timing and how to test apps on it.
Published on:
A security team asks an AI model to reproduce a known vulnerability so they can prove it is real and patch it, and the model refuses. According to Mistral's launch post, that refusal is why several leading closed models score near zero on one of the Artificial Analysis Cyber Index tests, which asks a model to reproduce a real vulnerability in open-source software and then patch it. Mistral Large 4 scores 82% on that test, the highest of any model.
Mistral Large 4, nicknamed Le Chonk, launched in public preview on October 6, 2026, with open weights promised for the end of the month. This guide covers its specs, the benchmark numbers Mistral published, what the open weights change for teams that self-host, and how to test a chat or tool-using agent with TestMu AI before you move it onto the new model.
TL;DR
Mistral Large 4, nicknamed Le Chonk, is Mistral AI's open-weight, natively multimodal mixture-of-experts model, launched in public preview on October 6, 2026. Mistral's launch post describes about 1 trillion parameters with 49 billion active, and the model is available through a preview API now, with open weights scheduled for the end of October.
- Model size: Mistral's launch post lists 1 trillion parameters with 49 billion active, while the Mistral Large 4 model card lists 1.05 trillion total, 52 billion active and a 1.6 billion-parameter vision encoder.
- Cybersecurity results: Mistral reports that Mistral Large 4 scores 82% on the Cyber Index test that reproduces and patches a real vulnerability (CyberGym-E2E in its charts), solves 93% of Cybench, and ranks in the index's top five.
- Open weights: Yes. Mistral says the Mistral Large 4 weights will be released at the end of October 2026, and the model runs as a public preview API on Mistral Studio until then.
- Context window: The Mistral Large 4 model card lists a 1M-token context window, with function calling, structured outputs and agents APIs.
- Testing apps built on it: TestMu AI Agent Testing scores and red-teams chat and voice agents built on Mistral Large 4, and TestMu AI Agent Assurance checks what tool-using agents built on it actually changed.
What Is Mistral Large 4?
Mistral's launch post calls Mistral Large 4 the company's largest and most capable model to date: an open-weight, natively multimodal mixture-of-experts model.
Mistral's own sources describe its size two ways, so the table below gives both, with the Mistral Large 4 model card as the source for the technical details:
| Spec | Mistral Large 4 | Source |
|---|---|---|
| Total parameters | 1 trillion (launch post); 1.05 trillion (model card) | Launch post, model card |
| Active parameters | 49 billion (launch post); 52 billion (model card) | Launch post, model card |
| Vision encoder | 1.6 billion parameters | Model card |
| Context window | 1M tokens | Model card |
| API features | Function calling, structured outputs, document Q&A, agents and conversations, built-in tools, batching | Model card |
| Training | Trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own datacenters in Europe | Launch post |
| Languages | Training data spans more than 160 languages, including every official EU language | Launch post |
| Availability | Public preview API on Mistral Studio as mistral-large-4; open weights at the end of October 2026 | Launch post, model card |
Mistral says it will share more on the architecture and its post-training method before the weights ship, so treat the preview specs as subject to change.
How Does Mistral Large 4 Perform on Benchmarks?
Mistral claims the model is the strongest open-weight model built in the US or Europe on aggregated benchmarks and state of the art among open models on cybersecurity, finance and legal work. These are the scores Mistral published for the preview:
| Area | Benchmark | Mistral Large 4 (as reported by Mistral) |
|---|---|---|
| Cybersecurity | CyberGym-E2E (reproduce and patch a real vulnerability) | 82%, the highest of any model |
| Cybersecurity | Cybench (40 challenges) | 93% solved |
| Cybersecurity | Artificial Analysis Cyber Index | Top five globally |
| Agentic coding | DeepSWE v1.1 | 61.7% |
| Agentic coding | SWE-Atlas-QnA | 59.4% |
| Agentic coding | Terminal-Bench 4 | 28.3% |
| Agentic coding | Coding Agent Index | 49.8% |
| Agentic workflows | AutomationBench (657 business workflows) | 59.9% |
| Knowledge work | AA-Briefcase | 1,393 Elo |
| Vision | Dense 200 (visual grounding) | 42%, against 41% for GPT-6-Astra |
| Human evaluation | Blind coding-quality ratings with Surge AI | 3.74 of 5, second of five models behind Claude Opus 5 (4.22) |
Read these as vendor numbers on a preview model. In its coverage of the launch, CNBC notes that the model still lags the frontier in areas such as coding.
The LLM benchmarks guide covers what each kind of benchmark can and cannot tell you about your own workload.
Why Is Mistral Large 4 Pitched for Cyber Defense?
Mistral's argument is about what the model will do as much as how well it scores. The launch post says closed models including Claude Opus 5.5 and GPT-6 Astra score near zero on that reproduce-and-patch test because they refuse the task, even though defending software often starts with proving a flaw is real. Mistral lists what the model handled in its own testing:
- Security operations work - analysing malware, prioritising vulnerabilities and writing detection rules, beyond what it was explicitly trained for.
- Self-deployment - running on private cloud or on-premise, so security teams can use it under their own policies without a provider-level refusal mid-incident.
- Pre-release red-teaming - until the weights ship, cybersecurity leaders, vetted partners and state authorities are testing the same model with reduced moderation and expanded cyber capabilities.
Co-founder and chief scientist Guillaume Lample told CNBC that the cyber defense capabilities "will enable enterprises and governments to defend themselves against threat actors that are jailbreaking closed models to perform cyber attacks." CNBC links the demand to the OpenAI Hugging Face incident, where Hugging Face's defense was blocked by closed models' guardrails and it turned to an open model instead.
When Will Mistral Large 4's Open Weights Be Released?
Mistral says the weights drop at the end of October 2026, and its announcement on X repeats the date. Until then the model is a public preview API, served from Mistral's European infrastructure. Open weights change who controls the model:
- Hosting - you can run the model in your own cloud or datacenter instead of calling a provider's API.
- Modification - CNBC notes that open models can be modified and self-hosted, in contrast to the leading closed systems from OpenAI and Anthropic.
- Sovereignty - Mistral also runs a European deployment end to end, independently of other digital service providers and under European law.
- Guardrails - when you host the weights, moderation is whatever you put around the model, so jailbreak and prompt-injection defenses become your team's job to build and test.
Mistral and TestMu AI are both signatories of the Open Weights and American AI Leadership letter; the open weights letter post explains why engineering teams need models they can inspect, run and adapt.
How Do You Test an App Built on Mistral Large 4?
Moving an agent onto a new model is a release, even when no line of your code changes: the same prompt can produce different answers, refusals and tool calls. With TestMu AI, Agent Testing grades what a chat or voice agent says and Agent Assurance checks what a tool-using agent does, so each risk below has a matching check:
| Risk when you adopt Mistral Large 4 | Check with | What the check looks at |
|---|---|---|
| Answers drift after the model swap | Agent Testing | The same regression scenarios, scored on the same metrics before and after the switch |
| A chat agent on self-hosted weights falls to a jailbreak or prompt injection | Agent Testing | AI red teaming across prompt injection, jailbreak and data exfiltration attacks |
| Wrong figures from a finance or legal assistant | Agent Testing | Hallucination scoring, plus Data Validation against your system of record |
| Uneven quality across languages | Agent Testing | Scenarios generated in the languages your users speak |
| A tool-using agent takes the wrong action | Agent Assurance | What the run changed, confirmed through a read-only check, not the agent's own summary |
Agent Testing: Grade What the Agent Says
Agent Testing uses AI testing agents to hold multi-turn conversations with your chat, voice or phone agent and score every reply, before and after deployment.
- Regression after a model update - scenarios carry a version history, so you can rerun the same versions when the underlying model changes and see drift as a score change instead of a customer complaint.
- Standard metrics - 9 metrics for chat and voice, including hallucination, bias, completeness and context awareness.
- AI red teaming - AI red teaming runs prompt injection, jailbreak and data exfiltration attacks across 9 attack categories at 3 intensity levels and grades each category from A+ to F.
- Multilingual scenarios - scenario generation, execution and evaluation are multilingual, which matters for a model trained on more than 160 languages.
Run the same suites from your terminal and CI with the Agent Testing CLI, and see agent regression testing for how to structure a before-and-after comparison.
Agent Assurance and rook: Grade What the Agent Did
Mistral pitches the model at agentic workflows, and the model card lists function calling, agents and built-in tools. Agent Assurance tests agents like these: its rook CLI reads your codebase (or, for an agent you can only reach through an API, works from the endpoint and spec you give it), writes functional, non-functional and adversarial scenarios, invokes the agent for real, and grades each criterion against what the run changed rather than what the agent reported.
- Evidence, not replies - tool calls are checked against the agent's own tool surface, and records are confirmed through a read-only query tool you approve.
- Three verdicts - every criterion is Pass, Fail or Unable to Verify, and Unable to Verify, which includes anything an API-only agent does not expose, is reported apart from the pass rate.
- Adversarial by default - prompt injection, jailbreak, data exfiltration, PII leakage and policy violation scenarios are generated as a class.
Agent Assurance is pre-alpha and publicly installable. Install the CLI and follow the Agent Assurance quickstart against a staging environment, since the agent's writes during a run are real:
npm install -g @testmuai/rookGetting Started With Mistral Large 4
Try the preview API on Mistral Studio with the model ID mistral-large-4, and estimate prompt sizes for your workload with TestMu AI's free token counter before you plan around the 1M-token window.
Before you switch a production agent, run your regression and red-team suites against both the current model and Mistral Large 4, and compare the results. If you plan to self-host when the weights ship, schedule the same runs against your own deployment, because your guardrails are now part of what you are testing.
Author
Anubhav Singhmaar is an AI Product Manager at TestMu AI driving Kane CLI, the command-line tool that brings browser automation to the terminal, turning natural-language flows into runs in a real Chrome browser that return pass or fail with shareable proof. He owns the roadmap and prioritization and works with engineering to ship developer-facing features. Before TestMu AI, he spent over four years at Sprinklr owning enterprise voice AI across APAC and EMEA. A mechanical engineer turned product manager, he grounds guidance in real QA workflows.
Reviewer
Sirajuddin Khan is Vice President of Product Management at TestMu AI (formerly LambdaTest), where he drives the company's agentic AI product strategy, building a suite of autonomous agents that includes Agentic Browsers and Agentic Visual Testing and shifting the unit of work from test execution to autonomous outcomes. One of the company's earliest product leaders, he has owned the roadmap for the high-performance execution cloud and grew the cross-browser testing products from early adoption to market leadership. He brings over a decade of experience across SaaS, B2B, and eCommerce, with earlier product roles at Wydr and ShopClues, where his catalog and search work cut delivery SLAs and lifted seller activity. Sirajuddin holds an MBA in Information Technology from Sikkim Manipal University and a B.Tech in Computer Science Engineering from Maharshi Dayanand University.
Mistral Large 4 FAQs
Did you find this page helpful?
More Related Blogs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests




