Why this benchmark is useful
Editorial brief pendingWe publish the methodology and ledger first; benchmark-specific analysis ships after desk review.
Benchmarks / Multimodal
MMMU
College-level questions that require reading images alongside text — charts, diagrams, chemical structures, medical scans — across 30 subjects.
Editorial brief pendingWe publish the methodology and ledger first; benchmark-specific analysis ships after desk review.
Reading guide pending. Use task format, scoring method, and source dates in the ledger until the desk brief ships.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Gemini 3.1 ProGoogle | 2026-02-19 | 72.4% | MMMU leaderboard2026-02-19 | Needs audit |
| 2 | GPT 5.5OpenAI | 2026-04-23 | 71.8% | MMMU leaderboard2026-04-23 | Needs audit |
| 3 | Claude Opus 4.8Anthropic | 2026-05-28 | 70.6% | MMMU leaderboard2026-05-28 | Needs audit |
| 4 | Gemini 3.5 FlashGoogle | 2026-05-19 | 68.9% | MMMU leaderboard2026-05-19 | Needs audit |
| 5 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 66.2% | MMMU leaderboard2026-02-17 | Needs audit |
| 6 | Llama 4 MaverickMeta | 2026-04-05 | 61.4% | MMMU leaderboard2026-04-05 | Needs audit |
| 7 | Gemini 1.0 UltraGoogle | 2024-02-08 | 59.4% | Gemini technical report2023-12-06 | Source-checked |
| 8 | GPT-4V (2023)OpenAI | 2023-09-25 | 56.8% | MMMU paper2023-11-27 | Source-checked |
Data-quality note pending. Every ledger row still carries source URL, source date, and ingest timestamp.
Human experts ~88%.
The Pack · Editorial newsletter
One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.