Why this benchmark is useful
Historical contest-math signal — once useful for hard integer reasoning. Solved for frontier models with extended thinking; keep for context, not current ranking.
AIME
Hard high-school competition math: 15 integer-answer problems per exam. Adopted as a frontier reasoning eval because each year's fresh paper resists contamination — until it leaks.
Historical contest-math signal — once useful for hard integer reasoning. Solved for frontier models with extended thinking; keep for context, not current ranking.
Do not use near-100% AIME headlines to separate today's frontier. Tiny yearly samples (15 problems) are noisy; always note exam year and attempt count.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | GPT-5 Codex (high)OpenAI | 2025-09-23 | 98.67% | 2026-09-01 | Source-checked |
| 2 | Gemini 3 Flash Preview (Reasoning)Google | 2025-12-17 | 97% | 2026-09-01 | Source-checked |
| 3 | GPT-5.2 (medium)OpenAI | 2025-12-11 | 96.67% | 2026-09-01 | Source-checked |
| 4 | MiMo-V2-Flash (Reasoning)Xiaomi | 2025-12-16 | 96.33% | 2026-09-01 | Source-checked |
| 5 | Gemini 3 ProGoogle | 2025-11-18 | 95.67% | 2026-09-01 | Source-checked |
| 6 | GPT-5.1 Codex (high)OpenAI | 2025-11-13 | 95.67% | 2026-09-01 | Source-checked |
| 7 | GLM-4.7 (Reasoning)Zhipu | 2025-12-22 | 95% | 2026-09-01 | Source-checked |
| 8 | Kimi K2 ThinkingMoonshot | 2025-11-06 | 94.67% | 2026-09-01 | Source-checked |
| 9 | GPT-5 (high)OpenAI | 2025-08-07 | 94.33% | 2026-09-01 | Source-checked |
| 1 | Claude Opus 4.8AnthropicAIME 2024 with extended thinking. | 2026-05-28 | 94.2% | Anthropic Claude 4 family2026-05-28 | Needs audit |
| 11 | GPT-5.1OpenAI | 2025-11-13 | 94% | 2026-09-01 | Source-checked |
| 2 | Claude 3.7 Sonnet (2025-02)AnthropicAIME 2024 with extended thinking enabled. | 2025-02-19 | 93.7% | Anthropic Claude 3.7 Sonnet announcement2025-02-19 | Source-checked |
| 13 | Grok 4xAI | 2025-07-10 | 92.67% | 2026-09-01 | Source-checked |
| 14 | DeepSeek V3.2 (Reasoning)DeepSeek | 2025-12-01 | 92% | 2026-09-01 | Source-checked |
| 3 | GPT 5.5OpenAI | 2026-04-23 | 91.8% | OpenAI model release notes2026-04-23 | Needs audit |
| 16 | GPT-5 (medium)OpenAI | 2025-08-07 | 91.67% | 2026-09-01 | Source-checked |
| 17 | Claude Opus 4.5 (Reasoning)Anthropic | 2025-11-24 | 91.33% | 2026-09-01 | Source-checked |
| 4 | GPT 5OpenAI | 2025-12-11 | 89.2% | OpenAI model release notes2025-12-01 | Needs audit |
| 5 | DeepSeek R1DeepSeekAIME 2024 with chain-of-thought. | 2025-01-20 | 88.4% | DeepSeek R1 model card2025-01-20 | Needs audit |
| 20 | OpenAI o3OpenAI | 2025-04-16 | 88.33% | 2026-09-01 | Source-checked |
| 21 | Claude 4.5 Sonnet (Reasoning)Anthropic | 2025-09-29 | 88% | 2026-09-01 | Source-checked |
| 6 | Gemini 3.1 ProGoogle | 2026-02-19 | 87.6% | Google Gemini 3.1 Pro2026-02-19 | Needs audit |
| 7 | o3 (cons@64, AIME 2025)OpenAIAIME 2025 paper; 64-sample consensus. Single-shot is markedly lower. | 2025-04-16 | 87.5% | OpenAI: Learning to Reason2025-01-20 | Source-checked |
| 24 | Gemini 3 Pro Preview (low)Google | 2025-11-18 | 86.67% | 2026-09-01 | Source-checked |
| 9 | Kimi K2.6Moonshot | 2026-04-20 | 85.2% | Moonshot Kimi K2.62026-04-20 | Needs audit |
| 10 | OpenAI o1 (AIME 2024, cons@64)OpenAIWith 64-sample consensus; single-shot is markedly lower. | 2024-09-12 | 83.3% | OpenAI: Learning to Reason2024-09-12 | Source-checked |
| 27 | GPT-5 (low)OpenAI | 2025-08-07 | 83% | 2026-09-01 | Source-checked |
| 28 | MiniMax-M2.1MiniMax | 2025-12-23 | 82.67% | 2026-09-01 | Source-checked |
| 29 | Claude 4.1 Opus (Reasoning)Anthropic | 2025-08-05 | 80.33% | 2026-09-01 | Source-checked |
| 30 | Claude 4 Opus (Reasoning)Anthropic | 2025-05-22 | 73.33% | 2026-09-01 | Source-checked |
| 31 | Claude Opus 4.5Anthropic | 2025-11-24 | 62.67% | 2026-09-01 | Source-checked |
| 32 | Gemini 3 Flash PreviewGoogle | 2025-12-17 | 55.67% | 2026-09-01 | Source-checked |
| 33 | Mistral Large 3Mistral | 2025-12-02 | 38% | 2026-09-01 | Source-checked |
| 34 | Llama 4 MaverickMeta | 2026-04-05 | 19.33% | 2026-09-01 | Source-checked |
| 35 | Llama 4 ScoutMeta | 2026-04-05 | 14% | 2026-09-01 | Source-checked |
| 36 | Command ACohere | 2026-03-15 | 13% | 2026-09-01 | Source-checked |
| 37 | Claude 3 OpusAnthropic | 2024-03-04 | 3.33% | 2026-09-01 | Source-checked |
Archived — GPT-5.2 and peers now report 100% on recent AIME papers with extended thinking. Treat as solved, not a frontier separator. Prefer /benchmarks/frontiermath.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
Hard high-school competition math: 15 integer-answer problems per exam. Adopted as a frontier reasoning eval because each year's fresh paper resists contamination — until it leaks.
A strong AIME result says nothing about:
Accuracy (often pass@1 or cons@64). Task format: 15 problems, integer answers 0-999; reported as accuracy on AIME 2024/2025 papers.
AIME is currently marked Defunct in the atlas.
Contamination risk for AIME is graded Medium contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.