Why this benchmark is useful
Historical broad knowledge exam across 57 subjects — useful glossary context for how models were once compared. Saturated at the frontier; not for current ranking.
MMLU
Broad multiple-choice knowledge across 57 subjects — from US history to college physics to law — answered in a four-option exam format.
Historical broad knowledge exam across 57 subjects — useful glossary context for how models were once compared. Saturated at the frontier; not for current ranking.
Plain accuracy on four-option MCQs. Frontier models cluster near the expert ceiling (~89.8%); treat high scores as table stakes, not a separator.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.8Anthropic5-shot CoT. MMLU is saturated at the frontier. | 2026-05-28 | 91.4% | Anthropic Claude 4 family2026-05-28 | Needs audit |
| 2 | Claude Opus 4.6AnthropicReported with 5-shot chain-of-thought prompting. | 2026-02-05 | 91.2% | Anthropic Claude 3.7/4.6 model card2026-01-01 | Source-checked |
| 3 | GPT 5.5OpenAIMMLU is saturated at the frontier — high scores are table stakes, not differentiators. | 2026-04-23 | 90.8% | OpenAI model release notes2025-12-01 | Needs audit |
| 4 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 90.6% | Anthropic Claude 3.7/4.6 model card2026-02-17 | Source-checked |
| 5 | Gemini 3.1 ProGoogle | 2026-02-19 | 90.3% | Google Gemini 3.1 Pro2026-02-19 | Needs audit |
| 6 | Gemini 1.5 Ultra (2024)GoogleReported with CoT@32 voting, not single-shot. | 2024-02-15 | 90% | Gemini technical report2023-12-06 | Source-checked |
| 7 | GPT 5.4OpenAI | 2026-03-05 | 89.8% | OpenAI GPT-5.4 release2026-03-05 | Needs audit |
| 8 | DeepSeek V4 Pro MaxDeepSeek | 2026-04-24 | 89.7% | DeepSeek V4 release2026-04-24 | Needs audit |
| 9 | Kimi K2.6Moonshot | 2026-04-20 | 89.1% | Moonshot Kimi K2.62026-04-20 | Needs audit |
| 10 | Qwen 3.7 MaxAlibaba | 2026-05-20 | 88.9% | Alibaba Qwen 3.7 Max2026-05-20 | Needs audit |
| 11 | Grok 4.3xAI | 2026-04-30 | 88.5% | xAI Grok 4.32026-04-30 | Needs audit |
| 12 | Gemini 3.5 FlashGoogle | 2026-05-19 | 88.4% | Google Gemini 3.5 Flash2026-05-19 | Needs audit |
| 13 | Mistral Large 3Mistral | 2025-12-02 | 87.8% | Mistral Large 3 model card2025-12-02 | Needs audit |
| 14 | DeepSeek R1DeepSeek | 2025-01-20 | 87.4% | DeepSeek R1 model card2025-01-20 | Needs audit |
| 15 | Gemini 3 FlashGoogle | 2025-12-17 | 87.2% | Google Gemini 3 Flash2025-12-17 | Needs audit |
| 16 | Claude 3 OpusAnthropic | 2024-03-04 | 86.8% | Anthropic Claude 3 model card2024-03-04 | Source-checked |
| 17 | GPT-4 (2023)OpenAI | 2023-03-14 | 86.4% | GPT-4 Technical Report2023-03-15 | Source-checked |
| 18 | Kimi K2.5Moonshot | 2026-01-27 | 86.4% | Moonshot Kimi K2.52026-01-27 | Needs audit |
| 19 | Llama 4 MaverickMeta | 2026-04-05 | 86.1% | Meta Llama 4 model card2026-04-05 | Needs audit |
| 20 | Claude Haiku 4.5Anthropic | 2025-10-01 | 85.6% | Anthropic Claude Haiku 4.52025-10-01 | Needs audit |
| 21 | Llama 4 ScoutMeta | 2026-04-05 | 85.2% | Meta Llama 4 Scout model card2026-04-05 | Needs audit |
| 22 | MiniMax M3MiniMax | 2026-05-31 | 84.8% | MiniMax M32026-05-31 | Needs audit |
Saturated glossary entry — useful for context, not frontier discrimination. Inline vendor rows may lag the ledger; treat ceiling clustering as the honest read.
Expert-level humans ~89.8%; frontier models now cluster near the ceiling, so MMLU no longer separates the best systems.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
Broad multiple-choice knowledge across 57 subjects — from US history to college physics to law — answered in a four-option exam format.
A strong MMLU result says nothing about:
Accuracy (% correct), often few-shot. Task format: ~14,000 four-option multiple-choice questions; score is plain accuracy.
MMLU is currently marked Saturated in the atlas. Ceiling context: Expert-level humans ~89.8%; frontier models now cluster near the ceiling, so MMLU no longer separates the best systems.
Contamination risk for MMLU is graded High contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.