Why this benchmark is useful
When you need college-level multimodal understanding — charts, diagrams, structures, scans — not text-only knowledge exams.
MMMU
College-level questions that require reading images alongside text — charts, diagrams, chemical structures, medical scans — across 30 subjects.
When you need college-level multimodal understanding — charts, diagrams, structures, scans — not text-only knowledge exams.
Accuracy on image+text items. Frontier rows now sit at or above the ~88% human-expert ceiling, so this page is saturated — read MMMU-Pro for the live multimodal exam.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Gemini 3.1 ProGoogle | 2026-02-19 | 72.4% | MMMU leaderboard2026-02-19 | Needs audit |
| 2 | GPT 5.5OpenAI | 2026-04-23 | 71.8% | MMMU leaderboard2026-04-23 | Needs audit |
| 3 | Claude Opus 4.8Anthropic | 2026-05-28 | 70.6% | MMMU leaderboard2026-05-28 | Needs audit |
| 4 | Gemini 3.5 FlashGoogle | 2026-05-19 | 68.9% | MMMU leaderboard2026-05-19 | Needs audit |
| 5 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 66.2% | MMMU leaderboard2026-02-17 | Needs audit |
| 6 | Llama 4 MaverickMeta | 2026-04-05 | 61.4% | MMMU leaderboard2026-04-05 | Needs audit |
| 7 | Gemini 1.0 UltraGoogle | 2024-02-08 | 59.4% | Gemini technical report2023-12-06 | Source-checked |
| 8 | GPT-4V (2023)OpenAI | 2023-09-25 | 56.8% | MMMU paper2023-11-27 | Source-checked |
This benchmark sits in the multimodal family. It uses ~11,500 multimodal questions (image + text), multiple-choice and open-ended. The score should travel with its task format, scoring method, source date, and benchmark version.
Human experts ~88%.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
College-level questions that require reading images alongside text — charts, diagrams, chemical structures, medical scans — across 30 subjects.
A strong MMMU result says nothing about:
Accuracy (% correct). Task format: ~11,500 multimodal questions (image + text), multiple-choice and open-ended.
MMMU is currently marked Saturated in the atlas. Ceiling context: Human experts ~88%.
Contamination risk for MMMU is graded Medium contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.