Why this benchmark is useful
It gives a crisp reasoning signal with right-or-wrong answers. Use it to compare structured problem solving, while watching sample size and pass@k settings.
Benchmarks / Math
AIME
Hard high-school competition math: 15 integer-answer problems per exam. Adopted as a frontier reasoning eval because each year's fresh paper resists contamination — until it leaks.
It gives a crisp reasoning signal with right-or-wrong answers. Use it to compare structured problem solving, while watching sample size and pass@k settings.
Read AIME as a defunct signal with medium contamination risk. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Gemini 3 ProGoogle | 2025-11-18 | 95.67% | 2026-07-21 | Source-checked |
| 1 | Claude Opus 4.8AnthropicAIME 2024 with extended thinking. | 2026-05-28 | 94.2% | Anthropic Claude 4 family2026-05-28 | Needs audit |
| 3 | GPT-5.1OpenAI | 2025-11-13 | 94% | 2026-07-21 | Source-checked |
| 2 | Claude 3.7 Sonnet (2025-02)AnthropicAIME 2024 with extended thinking enabled. | 2025-02-19 | 93.7% | Anthropic Claude 3.7 Sonnet announcement2025-02-19 | Source-checked |
| 3 | GPT 5.5OpenAI | 2026-04-23 | 91.8% | OpenAI model release notes2026-04-23 | Needs audit |
| 4 | GPT 5OpenAI | 2025-12-11 | 89.2% | OpenAI model release notes2025-12-01 | Needs audit |
| 5 | DeepSeek R1DeepSeekAIME 2024 with chain-of-thought. | 2025-01-20 | 88.4% | DeepSeek R1 model card2025-01-20 | Needs audit |
| 8 | OpenAI o3OpenAI | 2025-04-16 | 88.33% | 2026-07-21 | Source-checked |
| 6 | Gemini 3.1 ProGoogle | 2026-02-19 | 87.6% | Google Gemini 3.1 Pro2026-02-19 | Needs audit |
| 7 | o3 (cons@64, AIME 2025)OpenAIAIME 2025 paper; 64-sample consensus. Single-shot is markedly lower. | 2025-04-16 | 87.5% | OpenAI: Learning to Reason2025-01-20 | Source-checked |
| 9 | Kimi K2.6Moonshot | 2026-04-20 | 85.2% | Moonshot Kimi K2.62026-04-20 | Needs audit |
| 10 | OpenAI o1 (AIME 2024, cons@64)OpenAIWith 64-sample consensus; single-shot is markedly lower. | 2024-09-12 | 83.3% | OpenAI: Learning to Reason2024-09-12 | Source-checked |
| 13 | Claude Opus 4.5Anthropic | 2025-11-24 | 62.67% | 2026-07-21 | Source-checked |
| 14 | Gemini 3 Flash PreviewGoogle | 2025-12-17 | 55.67% | 2026-07-21 | Source-checked |
| 15 | Mistral Large 3Mistral | 2025-12-02 | 38% | 2026-07-21 | Source-checked |
| 16 | Llama 4 MaverickMeta | 2026-04-05 | 19.33% | 2026-07-21 | Source-checked |
| 17 | Llama 4 ScoutMeta | 2026-04-05 | 14% | 2026-07-21 | Source-checked |
| 18 | Command ACohere | 2026-03-15 | 13% | 2026-07-21 | Source-checked |
| 19 | Claude 3 OpusAnthropic | 2024-03-04 | 3.33% | 2026-07-21 | Source-checked |
Archived — GPT-5.2 and peers now report 100% on recent AIME papers with extended thinking. Treat as solved, not a frontier separator.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
The Pack · Editorial newsletter
One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.