Why this benchmark is useful
When classic MMLU is too easy or too guessable — a stricter ten-option, reasoning-heavy knowledge exam across domains.
A harder MMLU successor: ten answer options instead of four, reasoning-heavy questions, and a scrubbed item set built to be less guessable.
When classic MMLU is too easy or too guessable — a stricter ten-option, reasoning-heavy knowledge exam across domains.
Accuracy (% correct), usually with chain-of-thought. Still multiple-choice; prompt style (CoT vs direct) can move scores several points.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Gemini 3 ProGoogle | 2025-11-18 | 89.8% | 2026-09-01 | Source-checked |
| 2 | Claude Opus 4.5 (Reasoning)Anthropic | 2025-11-24 | 89.5% | 2026-09-01 | Source-checked |
| 3 | Gemini 3 Pro Preview (low)Google | 2025-11-18 | 89.5% | 2026-09-01 | Source-checked |
| 4 | Gemini 3 Flash Preview (Reasoning)Google | 2025-12-17 | 89% | 2026-09-01 | Source-checked |
| 5 | Claude Opus 4.5Anthropic | 2025-11-24 | 88.9% | 2026-09-01 | Source-checked |
| 6 | Gemini 3 Flash PreviewGoogle | 2025-12-17 | 88.2% | 2026-09-01 | Source-checked |
| 7 | Claude 4.1 Opus (Reasoning)Anthropic | 2025-08-05 | 88% | 2026-09-01 | Source-checked |
| 8 | Claude 4.5 Sonnet (Reasoning)Anthropic | 2025-09-29 | 87.5% | 2026-09-01 | Source-checked |
| 9 | MiniMax-M2.1MiniMax | 2025-12-23 | 87.5% | 2026-09-01 | Source-checked |
| 10 | GPT-5.2OpenAI | 2025-12-11 | 87.4% | 2026-09-01 | Source-checked |
| 11 | Claude 4 Opus (Reasoning)Anthropic | 2025-05-22 | 87.3% | 2026-09-01 | Source-checked |
| 12 | GPT-5 (high)OpenAI | 2025-08-07 | 87.1% | 2026-09-01 | Source-checked |
| 13 | GPT-5.1OpenAI | 2025-11-13 | 87% | 2026-09-01 | Source-checked |
| 14 | GPT-5 (medium)OpenAI | 2025-08-07 | 86.7% | 2026-09-01 | Source-checked |
| 15 | Grok 4xAI | 2025-07-10 | 86.6% | 2026-09-01 | Source-checked |
| 16 | GPT-5 Codex (high)OpenAI | 2025-09-23 | 86.5% | 2026-09-01 | Source-checked |
| 17 | DeepSeek V3.2 (Reasoning)DeepSeek | 2025-12-01 | 86.2% | 2026-09-01 | Source-checked |
| 18 | GPT-5 (low)OpenAI | 2025-08-07 | 86% | 2026-09-01 | Source-checked |
| 19 | GPT-5.1 Codex (high)OpenAI | 2025-11-13 | 86% | 2026-09-01 | Source-checked |
| 20 | GPT-5.2 (medium)OpenAI | 2025-12-11 | 85.9% | 2026-09-01 | Source-checked |
| 21 | GLM-4.7 (Reasoning)Zhipu | 2025-12-22 | 85.6% | 2026-09-01 | Source-checked |
| 22 | OpenAI o3OpenAI | 2025-04-16 | 85.3% | 2026-09-01 | Source-checked |
| 23 | DeepSeek R1DeepSeek | 2025-01-20 | 84.9% | 2026-09-01 | Source-checked |
| 24 | Kimi K2 ThinkingMoonshot | 2025-11-06 | 84.8% | 2026-09-01 | Source-checked |
| 25 | MiMo-V2-Flash (Reasoning)Xiaomi | 2025-12-16 | 84.3% | 2026-09-01 | Source-checked |
| 26 | OpenAI o1OpenAI | 2024-09-12 | 84.1% | 2026-09-01 | Source-checked |
| 1 | Claude Opus 4.8AnthropicCoT prompting. Harder than base MMLU — gaps are wider at the frontier. | 2026-05-28 | 82.4% | MMLU-Pro HuggingFace leaderboard2026-05-28 | Needs audit |
| 2 | GPT 5.5OpenAI | 2026-04-23 | 81.6% | MMLU-Pro HuggingFace leaderboard2026-04-23 | Needs audit |
| 29 | Llama 4 MaverickMeta | 2026-04-05 | 80.9% | 2026-09-01 | Source-checked |
| 3 | Gemini 3.1 ProGoogle | 2026-02-19 | 79.8% | MMLU-Pro HuggingFace leaderboard2026-02-19 | Needs audit |
| 4 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 78.2% | MMLU-Pro HuggingFace leaderboard2026-02-17 | Needs audit |
| 5 | DeepSeek V4 Pro MaxDeepSeek | 2026-04-24 | 76.4% | MMLU-Pro HuggingFace leaderboard2026-04-24 | Needs audit |
| 33 | Llama 4 ScoutMeta | 2026-04-05 | 75.2% | 2026-09-01 | Source-checked |
| 6 | Qwen 3.7 MaxAlibaba | 2026-05-20 | 75.1% | MMLU-Pro HuggingFace leaderboard2026-05-20 | Needs audit |
| 7 | Kimi K2.6Moonshot | 2026-04-20 | 74.6% | MMLU-Pro HuggingFace leaderboard2026-04-20 | Needs audit |
| 8 | Mistral Large 3Mistral | 2025-12-02 | 72.8% | MMLU-Pro HuggingFace leaderboard2025-12-02 | Needs audit |
| 9 | GPT-4o (2024)OpenAI | 2024-05-01 | 72.6% | MMLU-Pro paper2024-06-03 | Source-checked |
| 38 | Command ACohere | 2026-03-15 | 71.2% | 2026-09-01 | Source-checked |
| 10 | Claude 3 OpusAnthropic | 2024-03-04 | 68.4% | MMLU-Pro paper2024-06-03 | Source-checked |
This benchmark sits in the knowledge family. It uses ~12,000 ten-option questions across 14 domains; chain-of-thought is expected. The score should travel with its task format, scoring method, source date, and benchmark version.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
A harder MMLU successor: ten answer options instead of four, reasoning-heavy questions, and a scrubbed item set built to be less guessable.
A strong MMLU-Pro result says nothing about:
Accuracy (% correct) with CoT. Task format: ~12,000 ten-option questions across 14 domains; chain-of-thought is expected.
MMLU-Pro is currently marked Active in the atlas.
Contamination risk for MMLU-Pro is graded Medium contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.