Why this benchmark is useful
Original MMMU no longer separates the frontier. MMMU-Pro is the live multimodal exam: shortcuts removed, ten options, still a spread below the old 88% ceiling.
MMMU Pro
Harder multimodal college exam: original MMMU items with text-only shortcuts filtered out, ten answer options instead of four, and a vision-only setting where the question is inside the image. 3,460 questions across 30 subjects.
Original MMMU no longer separates the frontier. MMMU-Pro is the live multimodal exam: shortcuts removed, ten options, still a spread below the old 88% ceiling.
Read AA MMMU-Pro accuracy, not original MMMU or Vals 4-option headlines. Gemini 3.7 Flash leads the public AA board around 85%; Claude Opus 5 and GPT-5.6 Sol sit in the same band. Do not mix with Vals rows near 90%.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Gemini 3.7 FlashGoogle | 2026-08-13 | 85.5% | 2026-09-01 | Source-checked |
| 2 | Claude Opus 5Anthropic | 2026-07-24 | 84.7% | 2026-09-01 | Source-checked |
| 3 | Gemini 3.5 FlashGoogle | 2026-05-19 | 84.3% | 2026-09-01 | Source-checked |
| 4 | GPT-5.6 SolOpenAI | 2026-07-09 | 83.4% | 2026-09-01 | Source-checked |
| 5 | Gemini 3.6 FlashGoogle | 2026-07-21 | 83.2% | 2026-09-01 | Source-checked |
| 7 | Qwen3.8 MaxAlibaba | 2026-08-03 | 82.3% | 2026-09-01 | Source-checked |
| 8 | GPT-5.6 TerraOpenAI | 2026-07-09 | 80.7% | 2026-09-01 | Source-checked |
| 9 | Kimi K3Moonshot | 2026-07-16 | 80.5% | 2026-09-01 | Source-checked |
Ledger rows from Artificial Analysis' public MMMU-Pro evaluation (1 Sep 2026 board). The 2024 paper reported much lower scores on an earlier model set; 2026 frontier rows sit in the low-to-mid 80s. Vals runs a different, easier protocol under a similar name.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
Harder multimodal college exam: original MMMU items with text-only shortcuts filtered out, ten answer options instead of four, and a vision-only setting where the question is inside the image. 3,460 questions across 30 subjects.
A strong MMMU-Pro result says nothing about:
Accuracy (% correct). VerdictPal ledger rows are Artificial Analysis' independent MMMU-Pro run. Task format: 3,460 image+text (and vision-only) items across art, business, science, health, humanities, and engineering. Ten multiple-choice options after filtering questions a text-only model can solve.
MMMU-Pro is currently marked Active in the atlas.
Contamination risk for MMMU-Pro is graded Medium contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.