87.8%
MMLU
Models / Mistral
Mistral Large · Released 2025-12-02
Mistral's December 2025 flagship MoE. 262K context, vision input, $0.50/$1.50 per 1M tokens. 675B total / 41B active parameters.
European open-API frontier for multilingual enterprise chat, document QA, and vision+text workflows.
Family profile
Best published score in each covered benchmark family.
7 tested benchmark families
Where Mistral Large 3 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.
The three highest-scoring models with pages in each capability family. Where Mistral Large 3 shows up, it's highlighted.
Benchmark rows added to the public ledger for Mistral Large 3 in the last 120 days. Older rows live in the full table below.
| Benchmark | Family | Score | Source | Days ago |
|---|---|---|---|---|
| AA-LCR | Long context | 34.67% | Artificial Analysis | 5 |
| AA output speed | Performance | 78 t/s | Artificial Analysis | 5 |
| AA time to first token | Performance | 0.77s | Artificial Analysis | 5 |
| AIME | Math | 38% | Artificial Analysis | 5 |
| Artificial Analysis Intelligence Index | Reasoning | 15.9% | Artificial Analysis | 5 |
| Humanity's Last Exam | Reasoning | 4.2% | Artificial Analysis | 5 |
| IFBench | Reasoning | 36.19% | Artificial Analysis | 5 |
| LiveCodeBench | Coding | 46.5% | Artificial Analysis | 5 |
| SciCode | Coding | 36.2% | Artificial Analysis | 5 |
| τ³-Bench Banking | Agentic | 5.77% | Artificial Analysis | 5 |
| Terminal-Bench | Agentic | 11.99% | Artificial Analysis | 5 |
| SWE-bench Verified | Coding | 76.2% | SWE-bench Verified leaderboard | 89 |
| GPQA Diamond | Reasoning | 71.2% | Artificial Analysis GPQA Diamond evaluation | 89 |
| MCP Atlas | Tool use | 70.4% | MCP Atlas leaderboard | 89 |
A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.
Hover any dot for name, score, and input price. Keyboard: tab through the top twelve, or use the ranking below.
Every catalog benchmark for Mistral Large 3. Scores link to the original source; gaps mean no public row exists yet.
20 of 37 catalog benchmarks have a sourced row for Mistral Large 3.
15.9%
AA Intelligence Index v4.1
71.2%
GPQA Diamond
4.2%
HLE (AA run)
36.19%
IFBench (AA run)
62.8%
LiveBench
38%
AIME 2025 (AA run)
92.4% pass@1
HumanEval
46.5%
LiveCodeBench (AA run)
36.2%
SciCode (AA run)
76.2%
SWE-bench Verified
11.99%
Terminal-Bench v2.1 (AA run)
5.77%
τ³-Bench Banking (AA run)
34.67%
AA-LCR (AA run)
50.8%
LongBench v2
78 t/s
Median output tokens/s (1k prompt, default provider)
0.77s
Median time to first token (1k prompt, default provider)