Benchmarks / Multimodal

Massive Multi-discipline Multimodal Understanding

MMMU

College-level questions that require reading images alongside text — charts, diagrams, chemical structures, medical scans — across 30 subjects.

What this does not measure
  • Pure vision quality — many items can be partly solved from the text alone, inflating scores.
  • Real document workflows — it is an exam, not 'read this 40-page PDF and extract the table'.
Analysis

Why this benchmark is useful

When you need college-level multimodal understanding — charts, diagrams, structures, scans — not text-only knowledge exams.

Scope

Coverage map

Task family
Multimodal
Format
~11,500 multimodal questions (image + text), multiple-choice and open-ended.
Scoring
Accuracy (% correct).
Maintainer
Yue et al.
Reading guide

How to read the scores

Accuracy on image+text items. Frontier rows now sit at or above the ~88% human-expert ceiling, so this page is saturated — read MMMU-Pro for the live multimodal exam.

Blind spots

What it does not cover

  • Pure vision quality — many items can be partly solved from the text alone, inflating scores.
  • Real document workflows — it is an exam, not 'read this 40-page PDF and extract the table'.
Scores

Evidence ledger

8 rows
88 rows
72.4%best score
2source-checked
1sources
2026-05-28to 2023-11-27
Papers / model cards8

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Opus 4.870.6% · MMMU leaderboard
  2. Gemini 3.5 Flash68.9% · MMMU leaderboard
  3. GPT 5.571.8% · MMMU leaderboard
  4. Llama 4 Maverick61.4% · MMMU leaderboard
  5. Gemini 3.1 Pro72.4% · MMMU leaderboard

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Gemini 3.1 ProGoogle2026-02-1972.4%
MMMU leaderboard2026-02-19
Needs audit
2GPT 5.5OpenAI2026-04-2371.8%
MMMU leaderboard2026-04-23
Needs audit
3Claude Opus 4.8Anthropic2026-05-2870.6%
MMMU leaderboard2026-05-28
Needs audit
4Gemini 3.5 FlashGoogle2026-05-1968.9%
MMMU leaderboard2026-05-19
Needs audit
5Claude Sonnet 4.6Anthropic2026-02-1766.2%
MMMU leaderboard2026-02-17
Needs audit
6Llama 4 MaverickMeta2026-04-0561.4%
MMMU leaderboard2026-04-05
Needs audit
7Gemini 1.0 UltraGoogle2024-02-0859.4%
Gemini technical report2023-12-06
Source-checked
8GPT-4V (2023)OpenAI2023-09-2556.8%
MMMU paper2023-11-27
Source-checked
Method

What it covers

This benchmark sits in the multimodal family. It uses ~11,500 multimodal questions (image + text), multiple-choice and open-ended. The score should travel with its task format, scoring method, source date, and benchmark version.

Score ceiling

Where it breaks down

Human experts ~88%.

Tools that report it

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does MMMU measure?

College-level questions that require reading images alongside text — charts, diagrams, chemical structures, medical scans — across 30 subjects.

What does a high MMMU score not prove?

A strong MMMU result says nothing about:

  • Pure vision quality — many items can be partly solved from the text alone, inflating scores.
  • Real document workflows — it is an exam, not 'read this 40-page PDF and extract the table'.

How is MMMU scored?

Accuracy (% correct). Task format: ~11,500 multimodal questions (image + text), multiple-choice and open-ended.

Is MMMU saturated?

MMMU is currently marked Saturated in the atlas. Ceiling context: Human experts ~88%.

Can MMMU results be contaminated by training data?

Contamination risk for MMMU is graded Medium contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.