Why this benchmark is useful
Editorial brief pendingWe publish the methodology and ledger first; benchmark-specific analysis ships after desk review.
Benchmarks / Knowledge
A harder MMLU successor: ten answer options instead of four, reasoning-heavy questions, and a scrubbed item set built to be less guessable.
Editorial brief pendingWe publish the methodology and ledger first; benchmark-specific analysis ships after desk review.
Reading guide pending. Use task format, scoring method, and source dates in the ledger until the desk brief ships.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Gemini 3 ProGoogle | 2025-11-18 | 89.8% | 2026-07-21 | Source-checked |
| 2 | Claude Opus 4.5Anthropic | 2025-11-24 | 88.9% | 2026-07-21 | Source-checked |
| 3 | Gemini 3 Flash PreviewGoogle | 2025-12-17 | 88.2% | 2026-07-21 | Source-checked |
| 4 | GPT-5.2OpenAI | 2025-12-11 | 87.4% | 2026-07-21 | Source-checked |
| 5 | GPT-5.1OpenAI | 2025-11-13 | 87% | 2026-07-21 | Source-checked |
| 6 | OpenAI o3OpenAI | 2025-04-16 | 85.3% | 2026-07-21 | Source-checked |
| 7 | DeepSeek R1DeepSeek | 2025-01-20 | 84.9% | 2026-07-21 | Source-checked |
| 8 | OpenAI o1OpenAI | 2024-09-12 | 84.1% | 2026-07-21 | Source-checked |
| 1 | Claude Opus 4.8AnthropicCoT prompting. Harder than base MMLU — gaps are wider at the frontier. | 2026-05-28 | 82.4% | MMLU-Pro HuggingFace leaderboard2026-05-28 | Needs audit |
| 2 | GPT 5.5OpenAI | 2026-04-23 | 81.6% | MMLU-Pro HuggingFace leaderboard2026-04-23 | Needs audit |
| 11 | Llama 4 MaverickMeta | 2026-04-05 | 80.9% | 2026-07-21 | Source-checked |
| 3 | Gemini 3.1 ProGoogle | 2026-02-19 | 79.8% | MMLU-Pro HuggingFace leaderboard2026-02-19 | Needs audit |
| 4 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 78.2% | MMLU-Pro HuggingFace leaderboard2026-02-17 | Needs audit |
| 5 | DeepSeek V4 Pro MaxDeepSeek | 2026-04-24 | 76.4% | MMLU-Pro HuggingFace leaderboard2026-04-24 | Needs audit |
| 15 | Llama 4 ScoutMeta | 2026-04-05 | 75.2% | 2026-07-21 | Source-checked |
| 6 | Qwen 3.7 MaxAlibaba | 2026-05-20 | 75.1% | MMLU-Pro HuggingFace leaderboard2026-05-20 | Needs audit |
| 7 | Kimi K2.6Moonshot | 2026-04-20 | 74.6% | MMLU-Pro HuggingFace leaderboard2026-04-20 | Needs audit |
| 8 | Mistral Large 3Mistral | 2025-12-02 | 72.8% | MMLU-Pro HuggingFace leaderboard2025-12-02 | Needs audit |
| 9 | GPT-4o (2024)OpenAI | 2024-05-01 | 72.6% | MMLU-Pro paper2024-06-03 | Source-checked |
| 20 | Command ACohere | 2026-03-15 | 71.2% | 2026-07-21 | Source-checked |
| 10 | Claude 3 OpusAnthropic | 2024-03-04 | 68.4% | MMLU-Pro paper2024-06-03 | Source-checked |
Data-quality note pending. Every ledger row still carries source URL, source date, and ingest timestamp.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
The Pack · Editorial newsletter
One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.