Benchmarks / Knowledge

Massive Multitask Language Understanding

MMLU

Broad multiple-choice knowledge across 57 subjects — from US history to college physics to law — answered in a four-option exam format.

What this does not measure
  • Reasoning under uncertainty — a model can pattern-match the right letter without understanding the question.
  • Whether the model would give a safe, calibrated answer in the open-ended phrasing a real user types.
  • Anything past 2021 — the question bank is fixed and contains known errors in several subjects.
Analysis

Why this benchmark is useful

Historical broad knowledge exam across 57 subjects — useful glossary context for how models were once compared. Saturated at the frontier; not for current ranking.

Scope

Coverage map

Task family
Knowledge
Format
~14,000 four-option multiple-choice questions; score is plain accuracy.
Scoring
Accuracy (% correct), often few-shot.
Maintainer
Hendrycks et al.
Reading guide

How to read the scores

Plain accuracy on four-option MCQs. Frontier models cluster near the expert ceiling (~89.8%); treat high scores as table stakes, not a separator.

Blind spots

What it does not cover

  • Reasoning under uncertainty — a model can pattern-match the right letter without understanding the question.
  • Whether the model would give a safe, calibrated answer in the open-ended phrasing a real user types.
  • Anything past 2021 — the question bank is fixed and contains known errors in several subjects.
Scores

Evidence ledger

22 rows
2222 rows
91.4%best score
5source-checked
2sources
2026-05-31to 2023-03-15
Papers / model cards13

manual-snapshot · manual

Vendor claims9

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. MiniMax M384.8% · MiniMax M3
  2. Claude Opus 4.891.4% · Anthropic Claude 4 family
  3. Qwen 3.7 Max88.9% · Alibaba Qwen 3.7 Max
  4. Gemini 3.5 Flash88.4% · Google Gemini 3.5 Flash
  5. Grok 4.388.5% · xAI Grok 4.3

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Opus 4.8Anthropic5-shot CoT. MMLU is saturated at the frontier.2026-05-2891.4%
Anthropic Claude 4 family2026-05-28
Needs audit
2Claude Opus 4.6AnthropicReported with 5-shot chain-of-thought prompting.2026-02-0591.2%
Anthropic Claude 3.7/4.6 model card2026-01-01
Source-checked
3GPT 5.5OpenAIMMLU is saturated at the frontier — high scores are table stakes, not differentiators.2026-04-2390.8%
OpenAI model release notes2025-12-01
Needs audit
4Claude Sonnet 4.6Anthropic2026-02-1790.6%
Anthropic Claude 3.7/4.6 model card2026-02-17
Source-checked
5Gemini 3.1 ProGoogle2026-02-1990.3%
Google Gemini 3.1 Pro2026-02-19
Needs audit
6Gemini 1.5 Ultra (2024)GoogleReported with CoT@32 voting, not single-shot.2024-02-1590%
Gemini technical report2023-12-06
Source-checked
7GPT 5.4OpenAI2026-03-0589.8%
OpenAI GPT-5.4 release2026-03-05
Needs audit
8DeepSeek V4 Pro MaxDeepSeek2026-04-2489.7%
DeepSeek V4 release2026-04-24
Needs audit
9Kimi K2.6Moonshot2026-04-2089.1%
Moonshot Kimi K2.62026-04-20
Needs audit
10Qwen 3.7 MaxAlibaba2026-05-2088.9%
Alibaba Qwen 3.7 Max2026-05-20
Needs audit
11Grok 4.3xAI2026-04-3088.5%
xAI Grok 4.32026-04-30
Needs audit
12Gemini 3.5 FlashGoogle2026-05-1988.4%
Google Gemini 3.5 Flash2026-05-19
Needs audit
13Mistral Large 3Mistral2025-12-0287.8%
Mistral Large 3 model card2025-12-02
Needs audit
14DeepSeek R1DeepSeek2025-01-2087.4%
DeepSeek R1 model card2025-01-20
Needs audit
15Gemini 3 FlashGoogle2025-12-1787.2%
Google Gemini 3 Flash2025-12-17
Needs audit
16Claude 3 OpusAnthropic2024-03-0486.8%
Anthropic Claude 3 model card2024-03-04
Source-checked
17GPT-4 (2023)OpenAI2023-03-1486.4%
GPT-4 Technical Report2023-03-15
Source-checked
18Kimi K2.5Moonshot2026-01-2786.4%
Moonshot Kimi K2.52026-01-27
Needs audit
19Llama 4 MaverickMeta2026-04-0586.1%
Meta Llama 4 model card2026-04-05
Needs audit
20Claude Haiku 4.5Anthropic2025-10-0185.6%
Anthropic Claude Haiku 4.52025-10-01
Needs audit
21Llama 4 ScoutMeta2026-04-0585.2%
Meta Llama 4 Scout model card2026-04-05
Needs audit
22MiniMax M3MiniMax2026-05-3184.8%
MiniMax M32026-05-31
Needs audit
Method

What it covers

Saturated glossary entry — useful for context, not frontier discrimination. Inline vendor rows may lag the ledger; treat ceiling clustering as the honest read.

Score ceiling

Where it breaks down

Expert-level humans ~89.8%; frontier models now cluster near the ceiling, so MMLU no longer separates the best systems.

Tools that report it

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does MMLU measure?

Broad multiple-choice knowledge across 57 subjects — from US history to college physics to law — answered in a four-option exam format.

What does a high MMLU score not prove?

A strong MMLU result says nothing about:

  • Reasoning under uncertainty — a model can pattern-match the right letter without understanding the question.
  • Whether the model would give a safe, calibrated answer in the open-ended phrasing a real user types.
  • Anything past 2021 — the question bank is fixed and contains known errors in several subjects.

How is MMLU scored?

Accuracy (% correct), often few-shot. Task format: ~14,000 four-option multiple-choice questions; score is plain accuracy.

Is MMLU saturated?

MMLU is currently marked Saturated in the atlas. Ceiling context: Expert-level humans ~89.8%; frontier models now cluster near the ceiling, so MMLU no longer separates the best systems.

Can MMLU results be contaminated by training data?

Contamination risk for MMLU is graded High contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.