Benchmarks / Knowledge

Massive Multitask Language Understanding

MMLU

Broad multiple-choice knowledge across 57 subjects — from US history to college physics to law — answered in a four-option exam format.

KnowledgeSolidSaturatedHigh contamination riskSince 2021
What this does not measure
  • Reasoning under uncertainty — a model can pattern-match the right letter without understanding the question.
  • Whether the model would give a safe, calibrated answer in the open-ended phrasing a real user types.
  • Anything past 2021 — the question bank is fixed and contains known errors in several subjects.
Analysis

Why this benchmark is useful

It gives a compact signal for a specific capability. Use it as one dated receipt beside pricing, privacy, and hands-on evidence.

Scope

Coverage map

Task family
Knowledge
Format
~14,000 four-option multiple-choice questions; score is plain accuracy.
Scoring
Accuracy (% correct), often few-shot.
Maintainer
Hendrycks et al.
Reading guide

How to read the scores

Read MMLU as a saturated signal with high contamination risk. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.

Blind spots

What it does not cover

  • Reasoning under uncertainty — a model can pattern-match the right letter without understanding the question.
  • Whether the model would give a safe, calibrated answer in the open-ended phrasing a real user types.
  • Anything past 2021 — the question bank is fixed and contains known errors in several subjects.
Scores

Evidence ledger

22 rows
2222 rows
91.4%best score
5source-checked
2sources
2026-05-31to 2023-03-15
Papers / model cards13

manual-snapshot · manual

Vendor claims9

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. MiniMax M384.8% · MiniMax M3
  2. Claude Opus 4.891.4% · Anthropic Claude 4 family
  3. Qwen 3.7 Max88.9% · Alibaba Qwen 3.7 Max
  4. Gemini 3.5 Flash88.4% · Google Gemini 3.5 Flash
  5. Grok 4.388.5% · xAI Grok 4.3

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Opus 4.8Anthropic5-shot CoT. MMLU is saturated at the frontier.2026-05-2891.4%
Anthropic Claude 4 family2026-05-28
Needs audit
2Claude Opus 4.6AnthropicReported with 5-shot chain-of-thought prompting.2026-02-0591.2%
Anthropic Claude 3.7/4.6 model card2026-01-01
Source-checked
3GPT 5.5OpenAIMMLU is saturated at the frontier — high scores are table stakes, not differentiators.2026-04-2390.8%
OpenAI model release notes2025-12-01
Needs audit
4Claude Sonnet 4.6Anthropic2026-02-1790.6%
Anthropic Claude 3.7/4.6 model card2026-02-17
Source-checked
5Gemini 3.1 ProGoogle2026-02-1990.3%
Google Gemini 3.1 Pro2026-02-19
Needs audit
6Gemini 1.5 Ultra (2024)GoogleReported with CoT@32 voting, not single-shot.2024-02-1590%
Gemini technical report2023-12-06
Source-checked
7GPT 5.4OpenAI2026-03-0589.8%
OpenAI GPT-5.4 release2026-03-05
Needs audit
8DeepSeek V4 Pro MaxDeepSeek2026-04-2489.7%
DeepSeek V4 release2026-04-24
Needs audit
9Kimi K2.6Moonshot2026-04-2089.1%
Moonshot Kimi K2.62026-04-20
Needs audit
10Qwen 3.7 MaxAlibaba2026-05-2088.9%
Alibaba Qwen 3.7 Max2026-05-20
Needs audit
11Grok 4.3xAI2026-04-3088.5%
xAI Grok 4.32026-04-30
Needs audit
12Gemini 3.5 FlashGoogle2026-05-1988.4%
Google Gemini 3.5 Flash2026-05-19
Needs audit
13Mistral Large 3Mistral2025-12-0287.8%
Mistral Large 3 model card2025-12-02
Needs audit
14DeepSeek R1DeepSeek2025-01-2087.4%
DeepSeek R1 model card2025-01-20
Needs audit
15Gemini 3 FlashGoogle2025-12-1787.2%
Google Gemini 3 Flash2025-12-17
Needs audit
16Claude 3 OpusAnthropic2024-03-0486.8%
Anthropic Claude 3 model card2024-03-04
Source-checked
17GPT-4 (2023)OpenAI2023-03-1486.4%
GPT-4 Technical Report2023-03-15
Source-checked
18Kimi K2.5Moonshot2026-01-2786.4%
Moonshot Kimi K2.52026-01-27
Needs audit
19Llama 4 MaverickMeta2026-04-0586.1%
Meta Llama 4 model card2026-04-05
Needs audit
20Claude Haiku 4.5Anthropic2025-10-0185.6%
Anthropic Claude Haiku 4.52025-10-01
Needs audit
21Llama 4 ScoutMeta2026-04-0585.2%
Meta Llama 4 Scout model card2026-04-05
Needs audit
22MiniMax M3MiniMax2026-05-3184.8%
MiniMax M32026-05-31
Needs audit
Method

What it covers

Saturated glossary entry — useful for context, not frontier discrimination. Inline vendor rows may lag the ledger; treat ceiling clustering as the honest read.

Score ceiling

Where it breaks down

Expert-level humans ~89.8%; frontier models now cluster near the ceiling, so MMLU no longer separates the best systems.

Tools that report it

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.