Benchmarks / Multimodal

MMMU-Pro

MMMU Pro

Harder multimodal college exam: original MMMU items with text-only shortcuts filtered out, ten answer options instead of four, and a vision-only setting where the question is inside the image. 3,460 questions across 30 subjects.

What this does not measure
  • Original MMMU headlines — those still include text-solvable items. A 88%+ MMMU row is not an MMMU-Pro row.
  • Document workflows — still an exam, not 'read this 40-page PDF and extract the table'.
  • Vals' 'MMMU Pro' 4-option run — that board is near the 88.6% human-expert ceiling. VerdictPal rows on this page are Artificial Analysis' harder 10-option / vision-filter suite.
Analysis

Why this benchmark is useful

Original MMMU no longer separates the frontier. MMMU-Pro is the live multimodal exam: shortcuts removed, ten options, still a spread below the old 88% ceiling.

Scope

Coverage map

Task family
Multimodal
Format
3,460 image+text (and vision-only) items across art, business, science, health, humanities, and engineering. Ten multiple-choice options after filtering questions a text-only model can solve.
Scoring
Accuracy (% correct). VerdictPal ledger rows are Artificial Analysis' independent MMMU-Pro run.
Maintainer
Yue et al. / Artificial Analysis
Reading guide

How to read the scores

Read AA MMMU-Pro accuracy, not original MMMU or Vals 4-option headlines. Gemini 3.7 Flash leads the public AA board around 85%; Claude Opus 5 and GPT-5.6 Sol sit in the same band. Do not mix with Vals rows near 90%.

Blind spots

What it does not cover

  • Original MMMU headlines — those still include text-solvable items. A 88%+ MMMU row is not an MMMU-Pro row.
  • Document workflows — still an exam, not 'read this 40-page PDF and extract the table'.
  • Vals' 'MMMU Pro' 4-option run — that board is near the 88.6% human-expert ceiling. VerdictPal rows on this page are Artificial Analysis' harder 10-option / vision-filter suite.
Scores

Evidence ledger

8 rows
88 rows
85.5%best score
8source-checked
1sources
2026-09-01source date
8

api · api

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Gemini 3.7 Flash85.5% ·
  2. Claude Opus 584.7% ·
  3. Gemini 3.5 Flash84.3% ·
  4. GPT-5.6 Sol83.4% ·
  5. Gemini 3.6 Flash83.2% ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Gemini 3.7 FlashGoogle2026-08-1385.5%
2026-09-01
Source-checked
2Claude Opus 5Anthropic2026-07-2484.7%
2026-09-01
Source-checked
3Gemini 3.5 FlashGoogle2026-05-1984.3%
2026-09-01
Source-checked
4GPT-5.6 SolOpenAI2026-07-0983.4%
2026-09-01
Source-checked
5Gemini 3.6 FlashGoogle2026-07-2183.2%
2026-09-01
Source-checked
7Qwen3.8 MaxAlibaba2026-08-0382.3%
2026-09-01
Source-checked
8GPT-5.6 TerraOpenAI2026-07-0980.7%
2026-09-01
Source-checked
9Kimi K3Moonshot2026-07-1680.5%
2026-09-01
Source-checked
Method

What it covers

Ledger rows from Artificial Analysis' public MMMU-Pro evaluation (1 Sep 2026 board). The 2024 paper reported much lower scores on an earlier model set; 2026 frontier rows sit in the low-to-mid 80s. Vals runs a different, easier protocol under a similar name.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Tools that report it

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does MMMU-Pro measure?

Harder multimodal college exam: original MMMU items with text-only shortcuts filtered out, ten answer options instead of four, and a vision-only setting where the question is inside the image. 3,460 questions across 30 subjects.

What does a high MMMU-Pro score not prove?

A strong MMMU-Pro result says nothing about:

  • Original MMMU headlines — those still include text-solvable items. A 88%+ MMMU row is not an MMMU-Pro row.
  • Document workflows — still an exam, not 'read this 40-page PDF and extract the table'.
  • Vals' 'MMMU Pro' 4-option run — that board is near the 88.6% human-expert ceiling. VerdictPal rows on this page are Artificial Analysis' harder 10-option / vision-filter suite.

How is MMMU-Pro scored?

Accuracy (% correct). VerdictPal ledger rows are Artificial Analysis' independent MMMU-Pro run. Task format: 3,460 image+text (and vision-only) items across art, business, science, health, humanities, and engineering. Ten multiple-choice options after filtering questions a text-only model can solve.

Is MMMU-Pro saturated?

MMMU-Pro is currently marked Active in the atlas.

Can MMMU-Pro results be contaminated by training data?

Contamination risk for MMMU-Pro is graded Medium contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.