Benchmarks / Knowledge

MMLU-Pro

A harder MMLU successor: ten answer options instead of four, reasoning-heavy questions, and a scrubbed item set built to be less guessable.

KnowledgeSolidActiveMedium contamination riskSince 2024
What this does not measure
  • Real-world open-ended ability — it is still a multiple-choice exam, just a stricter one.
  • Robustness to prompt phrasing — scores swing several points with chain-of-thought vs direct answers.
Analysis

Why this benchmark is useful

Editorial brief pendingWe publish the methodology and ledger first; benchmark-specific analysis ships after desk review.

Scope

Coverage map

Task family
Knowledge
Format
~12,000 ten-option questions across 14 domains; chain-of-thought is expected.
Scoring
Accuracy (% correct) with CoT.
Maintainer
TIGER-Lab
Reading guide

How to read the scores

Reading guide pending. Use task format, scoring method, and source dates in the ledger until the desk brief ships.

Blind spots

What it does not cover

  • Real-world open-ended ability — it is still a multiple-choice exam, just a stricter one.
  • Robustness to prompt phrasing — scores swing several points with chain-of-thought vs direct answers.
Scores

Evidence ledger

21 rows
2121 rows
89.8%best score
13source-checked
2sources
2026-07-21to 2024-06-03
11

api · api

Papers / model cards10

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Gemini 3 Pro89.8% ·
  2. Claude Opus 4.588.9% ·
  3. Gemini 3 Flash Preview88.2% ·
  4. GPT-5.287.4% ·
  5. GPT-5.187% ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Gemini 3 ProGoogle2025-11-1889.8%
2026-07-21
Source-checked
2Claude Opus 4.5Anthropic2025-11-2488.9%
2026-07-21
Source-checked
3Gemini 3 Flash PreviewGoogle2025-12-1788.2%
2026-07-21
Source-checked
4GPT-5.2OpenAI2025-12-1187.4%
2026-07-21
Source-checked
5GPT-5.1OpenAI2025-11-1387%
2026-07-21
Source-checked
6OpenAI o3OpenAI2025-04-1685.3%
2026-07-21
Source-checked
7DeepSeek R1DeepSeek2025-01-2084.9%
2026-07-21
Source-checked
8OpenAI o1OpenAI2024-09-1284.1%
2026-07-21
Source-checked
1Claude Opus 4.8AnthropicCoT prompting. Harder than base MMLU — gaps are wider at the frontier.2026-05-2882.4%
MMLU-Pro HuggingFace leaderboard2026-05-28
Needs audit
2GPT 5.5OpenAI2026-04-2381.6%
MMLU-Pro HuggingFace leaderboard2026-04-23
Needs audit
11Llama 4 MaverickMeta2026-04-0580.9%
2026-07-21
Source-checked
3Gemini 3.1 ProGoogle2026-02-1979.8%
MMLU-Pro HuggingFace leaderboard2026-02-19
Needs audit
4Claude Sonnet 4.6Anthropic2026-02-1778.2%
MMLU-Pro HuggingFace leaderboard2026-02-17
Needs audit
5DeepSeek V4 Pro MaxDeepSeek2026-04-2476.4%
MMLU-Pro HuggingFace leaderboard2026-04-24
Needs audit
15Llama 4 ScoutMeta2026-04-0575.2%
2026-07-21
Source-checked
6Qwen 3.7 MaxAlibaba2026-05-2075.1%
MMLU-Pro HuggingFace leaderboard2026-05-20
Needs audit
7Kimi K2.6Moonshot2026-04-2074.6%
MMLU-Pro HuggingFace leaderboard2026-04-20
Needs audit
8Mistral Large 3Mistral2025-12-0272.8%
MMLU-Pro HuggingFace leaderboard2025-12-02
Needs audit
9GPT-4o (2024)OpenAI2024-05-0172.6%
MMLU-Pro paper2024-06-03
Source-checked
20Command ACohere2026-03-1571.2%
2026-07-21
Source-checked
10Claude 3 OpusAnthropic2024-03-0468.4%
MMLU-Pro paper2024-06-03
Source-checked
Method

What it covers

Data-quality note pending. Every ledger row still carries source URL, source date, and ingest timestamp.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Tools that report it

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.