Benchmarks / Knowledge

MMLU-Pro

A harder MMLU successor: ten answer options instead of four, reasoning-heavy questions, and a scrubbed item set built to be less guessable.

What this does not measure
  • Real-world open-ended ability — it is still a multiple-choice exam, just a stricter one.
  • Robustness to prompt phrasing — scores swing several points with chain-of-thought vs direct answers.
Analysis

Why this benchmark is useful

When classic MMLU is too easy or too guessable — a stricter ten-option, reasoning-heavy knowledge exam across domains.

Scope

Coverage map

Task family
Knowledge
Format
~12,000 ten-option questions across 14 domains; chain-of-thought is expected.
Scoring
Accuracy (% correct) with CoT.
Maintainer
TIGER-Lab
Reading guide

How to read the scores

Accuracy (% correct), usually with chain-of-thought. Still multiple-choice; prompt style (CoT vs direct) can move scores several points.

Blind spots

What it does not cover

  • Real-world open-ended ability — it is still a multiple-choice exam, just a stricter one.
  • Robustness to prompt phrasing — scores swing several points with chain-of-thought vs direct answers.
Scores

Evidence ledger

39 rows
3939 rows
89.8%best score
31source-checked
2sources
2026-09-01to 2024-06-03
29

api · api

Papers / model cards10

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Gemini 3 Pro89.8% ·
  2. Claude Opus 4.5 (Reasoning)89.5% ·
  3. Gemini 3 Pro Preview (low)89.5% ·
  4. Gemini 3 Flash Preview (Reasoning)89% ·
  5. Claude Opus 4.588.9% ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Gemini 3 ProGoogle2025-11-1889.8%
2026-09-01
Source-checked
2Claude Opus 4.5 (Reasoning)Anthropic2025-11-2489.5%
2026-09-01
Source-checked
3Gemini 3 Pro Preview (low)Google2025-11-1889.5%
2026-09-01
Source-checked
4Gemini 3 Flash Preview (Reasoning)Google2025-12-1789%
2026-09-01
Source-checked
5Claude Opus 4.5Anthropic2025-11-2488.9%
2026-09-01
Source-checked
6Gemini 3 Flash PreviewGoogle2025-12-1788.2%
2026-09-01
Source-checked
7Claude 4.1 Opus (Reasoning)Anthropic2025-08-0588%
2026-09-01
Source-checked
8Claude 4.5 Sonnet (Reasoning)Anthropic2025-09-2987.5%
2026-09-01
Source-checked
9MiniMax-M2.1MiniMax2025-12-2387.5%
2026-09-01
Source-checked
10GPT-5.2OpenAI2025-12-1187.4%
2026-09-01
Source-checked
11Claude 4 Opus (Reasoning)Anthropic2025-05-2287.3%
2026-09-01
Source-checked
12GPT-5 (high)OpenAI2025-08-0787.1%
2026-09-01
Source-checked
13GPT-5.1OpenAI2025-11-1387%
2026-09-01
Source-checked
14GPT-5 (medium)OpenAI2025-08-0786.7%
2026-09-01
Source-checked
15Grok 4xAI2025-07-1086.6%
2026-09-01
Source-checked
16GPT-5 Codex (high)OpenAI2025-09-2386.5%
2026-09-01
Source-checked
17DeepSeek V3.2 (Reasoning)DeepSeek2025-12-0186.2%
2026-09-01
Source-checked
18GPT-5 (low)OpenAI2025-08-0786%
2026-09-01
Source-checked
19GPT-5.1 Codex (high)OpenAI2025-11-1386%
2026-09-01
Source-checked
20GPT-5.2 (medium)OpenAI2025-12-1185.9%
2026-09-01
Source-checked
21GLM-4.7 (Reasoning)Zhipu2025-12-2285.6%
2026-09-01
Source-checked
22OpenAI o3OpenAI2025-04-1685.3%
2026-09-01
Source-checked
23DeepSeek R1DeepSeek2025-01-2084.9%
2026-09-01
Source-checked
24Kimi K2 ThinkingMoonshot2025-11-0684.8%
2026-09-01
Source-checked
25MiMo-V2-Flash (Reasoning)Xiaomi2025-12-1684.3%
2026-09-01
Source-checked
26OpenAI o1OpenAI2024-09-1284.1%
2026-09-01
Source-checked
1Claude Opus 4.8AnthropicCoT prompting. Harder than base MMLU — gaps are wider at the frontier.2026-05-2882.4%
MMLU-Pro HuggingFace leaderboard2026-05-28
Needs audit
2GPT 5.5OpenAI2026-04-2381.6%
MMLU-Pro HuggingFace leaderboard2026-04-23
Needs audit
29Llama 4 MaverickMeta2026-04-0580.9%
2026-09-01
Source-checked
3Gemini 3.1 ProGoogle2026-02-1979.8%
MMLU-Pro HuggingFace leaderboard2026-02-19
Needs audit
4Claude Sonnet 4.6Anthropic2026-02-1778.2%
MMLU-Pro HuggingFace leaderboard2026-02-17
Needs audit
5DeepSeek V4 Pro MaxDeepSeek2026-04-2476.4%
MMLU-Pro HuggingFace leaderboard2026-04-24
Needs audit
33Llama 4 ScoutMeta2026-04-0575.2%
2026-09-01
Source-checked
6Qwen 3.7 MaxAlibaba2026-05-2075.1%
MMLU-Pro HuggingFace leaderboard2026-05-20
Needs audit
7Kimi K2.6Moonshot2026-04-2074.6%
MMLU-Pro HuggingFace leaderboard2026-04-20
Needs audit
8Mistral Large 3Mistral2025-12-0272.8%
MMLU-Pro HuggingFace leaderboard2025-12-02
Needs audit
9GPT-4o (2024)OpenAI2024-05-0172.6%
MMLU-Pro paper2024-06-03
Source-checked
38Command ACohere2026-03-1571.2%
2026-09-01
Source-checked
10Claude 3 OpusAnthropic2024-03-0468.4%
MMLU-Pro paper2024-06-03
Source-checked
Method

What it covers

This benchmark sits in the knowledge family. It uses ~12,000 ten-option questions across 14 domains; chain-of-thought is expected. The score should travel with its task format, scoring method, source date, and benchmark version.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Tools that report it

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does MMLU-Pro measure?

A harder MMLU successor: ten answer options instead of four, reasoning-heavy questions, and a scrubbed item set built to be less guessable.

What does a high MMLU-Pro score not prove?

A strong MMLU-Pro result says nothing about:

  • Real-world open-ended ability — it is still a multiple-choice exam, just a stricter one.
  • Robustness to prompt phrasing — scores swing several points with chain-of-thought vs direct answers.

How is MMLU-Pro scored?

Accuracy (% correct) with CoT. Task format: ~12,000 ten-option questions across 14 domains; chain-of-thought is expected.

Is MMLU-Pro saturated?

MMLU-Pro is currently marked Active in the atlas.

Can MMLU-Pro results be contaminated by training data?

Contamination risk for MMLU-Pro is graded Medium contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.