Benchmarks / Multimodal

Massive Multi-discipline Multimodal Understanding

MMMU

College-level questions that require reading images alongside text — charts, diagrams, chemical structures, medical scans — across 30 subjects.

MultimodalSolidActiveMedium contamination riskSince 2023
What this does not measure
  • Pure vision quality — many items can be partly solved from the text alone, inflating scores.
  • Real document workflows — it is an exam, not 'read this 40-page PDF and extract the table'.
Analysis

Why this benchmark is useful

Editorial brief pendingWe publish the methodology and ledger first; benchmark-specific analysis ships after desk review.

Scope

Coverage map

Task family
Multimodal
Format
~11,500 multimodal questions (image + text), multiple-choice and open-ended.
Scoring
Accuracy (% correct).
Maintainer
Yue et al.
Reading guide

How to read the scores

Reading guide pending. Use task format, scoring method, and source dates in the ledger until the desk brief ships.

Blind spots

What it does not cover

  • Pure vision quality — many items can be partly solved from the text alone, inflating scores.
  • Real document workflows — it is an exam, not 'read this 40-page PDF and extract the table'.
Scores

Evidence ledger

8 rows
88 rows
72.4%best score
2source-checked
1sources
2026-05-28to 2023-11-27
Papers / model cards8

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Opus 4.870.6% · MMMU leaderboard
  2. Gemini 3.5 Flash68.9% · MMMU leaderboard
  3. GPT 5.571.8% · MMMU leaderboard
  4. Llama 4 Maverick61.4% · MMMU leaderboard
  5. Gemini 3.1 Pro72.4% · MMMU leaderboard

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Gemini 3.1 ProGoogle2026-02-1972.4%
MMMU leaderboard2026-02-19
Needs audit
2GPT 5.5OpenAI2026-04-2371.8%
MMMU leaderboard2026-04-23
Needs audit
3Claude Opus 4.8Anthropic2026-05-2870.6%
MMMU leaderboard2026-05-28
Needs audit
4Gemini 3.5 FlashGoogle2026-05-1968.9%
MMMU leaderboard2026-05-19
Needs audit
5Claude Sonnet 4.6Anthropic2026-02-1766.2%
MMMU leaderboard2026-02-17
Needs audit
6Llama 4 MaverickMeta2026-04-0561.4%
MMMU leaderboard2026-04-05
Needs audit
7Gemini 1.0 UltraGoogle2024-02-0859.4%
Gemini technical report2023-12-06
Source-checked
8GPT-4V (2023)OpenAI2023-09-2556.8%
MMMU paper2023-11-27
Source-checked
Method

What it covers

Data-quality note pending. Every ledger row still carries source URL, source date, and ingest timestamp.

Score ceiling

Where it breaks down

Human experts ~88%.

Tools that report it

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.