Why this benchmark is useful
It gives a compact signal for a specific capability. Use it as one dated receipt beside pricing, privacy, and hands-on evidence.
Benchmarks / Knowledge
MMLU
Broad multiple-choice knowledge across 57 subjects — from US history to college physics to law — answered in a four-option exam format.
It gives a compact signal for a specific capability. Use it as one dated receipt beside pricing, privacy, and hands-on evidence.
Read MMLU as a saturated signal with high contamination risk. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.8Anthropic5-shot CoT. MMLU is saturated at the frontier. | 2026-05-28 | 91.4% | Anthropic Claude 4 family2026-05-28 | Needs audit |
| 2 | Claude Opus 4.6AnthropicReported with 5-shot chain-of-thought prompting. | 2026-02-05 | 91.2% | Anthropic Claude 3.7/4.6 model card2026-01-01 | Source-checked |
| 3 | GPT 5.5OpenAIMMLU is saturated at the frontier — high scores are table stakes, not differentiators. | 2026-04-23 | 90.8% | OpenAI model release notes2025-12-01 | Needs audit |
| 4 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 90.6% | Anthropic Claude 3.7/4.6 model card2026-02-17 | Source-checked |
| 5 | Gemini 3.1 ProGoogle | 2026-02-19 | 90.3% | Google Gemini 3.1 Pro2026-02-19 | Needs audit |
| 6 | Gemini 1.5 Ultra (2024)GoogleReported with CoT@32 voting, not single-shot. | 2024-02-15 | 90% | Gemini technical report2023-12-06 | Source-checked |
| 7 | GPT 5.4OpenAI | 2026-03-05 | 89.8% | OpenAI GPT-5.4 release2026-03-05 | Needs audit |
| 8 | DeepSeek V4 Pro MaxDeepSeek | 2026-04-24 | 89.7% | DeepSeek V4 release2026-04-24 | Needs audit |
| 9 | Kimi K2.6Moonshot | 2026-04-20 | 89.1% | Moonshot Kimi K2.62026-04-20 | Needs audit |
| 10 | Qwen 3.7 MaxAlibaba | 2026-05-20 | 88.9% | Alibaba Qwen 3.7 Max2026-05-20 | Needs audit |
| 11 | Grok 4.3xAI | 2026-04-30 | 88.5% | xAI Grok 4.32026-04-30 | Needs audit |
| 12 | Gemini 3.5 FlashGoogle | 2026-05-19 | 88.4% | Google Gemini 3.5 Flash2026-05-19 | Needs audit |
| 13 | Mistral Large 3Mistral | 2025-12-02 | 87.8% | Mistral Large 3 model card2025-12-02 | Needs audit |
| 14 | DeepSeek R1DeepSeek | 2025-01-20 | 87.4% | DeepSeek R1 model card2025-01-20 | Needs audit |
| 15 | Gemini 3 FlashGoogle | 2025-12-17 | 87.2% | Google Gemini 3 Flash2025-12-17 | Needs audit |
| 16 | Claude 3 OpusAnthropic | 2024-03-04 | 86.8% | Anthropic Claude 3 model card2024-03-04 | Source-checked |
| 17 | GPT-4 (2023)OpenAI | 2023-03-14 | 86.4% | GPT-4 Technical Report2023-03-15 | Source-checked |
| 18 | Kimi K2.5Moonshot | 2026-01-27 | 86.4% | Moonshot Kimi K2.52026-01-27 | Needs audit |
| 19 | Llama 4 MaverickMeta | 2026-04-05 | 86.1% | Meta Llama 4 model card2026-04-05 | Needs audit |
| 20 | Claude Haiku 4.5Anthropic | 2025-10-01 | 85.6% | Anthropic Claude Haiku 4.52025-10-01 | Needs audit |
| 21 | Llama 4 ScoutMeta | 2026-04-05 | 85.2% | Meta Llama 4 Scout model card2026-04-05 | Needs audit |
| 22 | MiniMax M3MiniMax | 2026-05-31 | 84.8% | MiniMax M32026-05-31 | Needs audit |
Saturated glossary entry — useful for context, not frontier discrimination. Inline vendor rows may lag the ledger; treat ceiling clustering as the honest read.
Expert-level humans ~89.8%; frontier models now cluster near the ceiling, so MMLU no longer separates the best systems.
The Pack · Editorial newsletter
One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.