Benchmarks / Math

American Invitational Mathematics Examination

AIME

Hard high-school competition math: 15 integer-answer problems per exam. Adopted as a frontier reasoning eval because each year's fresh paper resists contamination — until it leaks.

What this does not measure
  • Tiny sample (15 problems/year) — one problem is ~6.7%, so single-run scores are noisy; demand pass@k or multiple seeds.
  • General capability — strong AIME scores correlate with reasoning training, not with everyday helpfulness.
  • Which year — '90% on AIME' is meaningless without the exam year and number of attempts.
Analysis

Why this benchmark is useful

Historical contest-math signal — once useful for hard integer reasoning. Solved for frontier models with extended thinking; keep for context, not current ranking.

Scope

Coverage map

Task family
Math
Format
15 problems, integer answers 0-999; reported as accuracy on AIME 2024/2025 papers.
Scoring
Accuracy (often pass@1 or cons@64).
Maintainer
MAA (exam) · used as an LLM eval
Reading guide

How to read the scores

Do not use near-100% AIME headlines to separate today's frontier. Tiny yearly samples (15 problems) are noisy; always note exam year and attempt count.

Blind spots

What it does not cover

  • Tiny sample (15 problems/year) — one problem is ~6.7%, so single-run scores are noisy; demand pass@k or multiple seeds.
  • General capability — strong AIME scores correlate with reasoning training, not with everyday helpfulness.
  • Which year — '90% on AIME' is meaningless without the exam year and number of attempts.
Scores

Evidence ledger

37 rows
3737 rows
98.67%best score
31source-checked
3sources
2026-09-01to 2024-09-12
28

api · api

Papers / model cards2

manual-snapshot · manual

Vendor claims7

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. GPT-5 Codex (high)98.67% ·
  2. Gemini 3 Flash Preview (Reasoning)97% ·
  3. GPT-5.2 (medium)96.67% ·
  4. MiMo-V2-Flash (Reasoning)96.33% ·
  5. Gemini 3 Pro95.67% ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1GPT-5 Codex (high)OpenAI2025-09-2398.67%
2026-09-01
Source-checked
2Gemini 3 Flash Preview (Reasoning)Google2025-12-1797%
2026-09-01
Source-checked
3GPT-5.2 (medium)OpenAI2025-12-1196.67%
2026-09-01
Source-checked
4MiMo-V2-Flash (Reasoning)Xiaomi2025-12-1696.33%
2026-09-01
Source-checked
5Gemini 3 ProGoogle2025-11-1895.67%
2026-09-01
Source-checked
6GPT-5.1 Codex (high)OpenAI2025-11-1395.67%
2026-09-01
Source-checked
7GLM-4.7 (Reasoning)Zhipu2025-12-2295%
2026-09-01
Source-checked
8Kimi K2 ThinkingMoonshot2025-11-0694.67%
2026-09-01
Source-checked
9GPT-5 (high)OpenAI2025-08-0794.33%
2026-09-01
Source-checked
1Claude Opus 4.8AnthropicAIME 2024 with extended thinking.2026-05-2894.2%
Anthropic Claude 4 family2026-05-28
Needs audit
11GPT-5.1OpenAI2025-11-1394%
2026-09-01
Source-checked
2Claude 3.7 Sonnet (2025-02)AnthropicAIME 2024 with extended thinking enabled.2025-02-1993.7%
Anthropic Claude 3.7 Sonnet announcement2025-02-19
Source-checked
13Grok 4xAI2025-07-1092.67%
2026-09-01
Source-checked
14DeepSeek V3.2 (Reasoning)DeepSeek2025-12-0192%
2026-09-01
Source-checked
3GPT 5.5OpenAI2026-04-2391.8%
OpenAI model release notes2026-04-23
Needs audit
16GPT-5 (medium)OpenAI2025-08-0791.67%
2026-09-01
Source-checked
17Claude Opus 4.5 (Reasoning)Anthropic2025-11-2491.33%
2026-09-01
Source-checked
4GPT 5OpenAI2025-12-1189.2%
OpenAI model release notes2025-12-01
Needs audit
5DeepSeek R1DeepSeekAIME 2024 with chain-of-thought.2025-01-2088.4%
DeepSeek R1 model card2025-01-20
Needs audit
20OpenAI o3OpenAI2025-04-1688.33%
2026-09-01
Source-checked
21Claude 4.5 Sonnet (Reasoning)Anthropic2025-09-2988%
2026-09-01
Source-checked
6Gemini 3.1 ProGoogle2026-02-1987.6%
Google Gemini 3.1 Pro2026-02-19
Needs audit
7o3 (cons@64, AIME 2025)OpenAIAIME 2025 paper; 64-sample consensus. Single-shot is markedly lower.2025-04-1687.5%
OpenAI: Learning to Reason2025-01-20
Source-checked
24Gemini 3 Pro Preview (low)Google2025-11-1886.67%
2026-09-01
Source-checked
9Kimi K2.6Moonshot2026-04-2085.2%
Moonshot Kimi K2.62026-04-20
Needs audit
10OpenAI o1 (AIME 2024, cons@64)OpenAIWith 64-sample consensus; single-shot is markedly lower.2024-09-1283.3%
OpenAI: Learning to Reason2024-09-12
Source-checked
27GPT-5 (low)OpenAI2025-08-0783%
2026-09-01
Source-checked
28MiniMax-M2.1MiniMax2025-12-2382.67%
2026-09-01
Source-checked
29Claude 4.1 Opus (Reasoning)Anthropic2025-08-0580.33%
2026-09-01
Source-checked
30Claude 4 Opus (Reasoning)Anthropic2025-05-2273.33%
2026-09-01
Source-checked
31Claude Opus 4.5Anthropic2025-11-2462.67%
2026-09-01
Source-checked
32Gemini 3 Flash PreviewGoogle2025-12-1755.67%
2026-09-01
Source-checked
33Mistral Large 3Mistral2025-12-0238%
2026-09-01
Source-checked
34Llama 4 MaverickMeta2026-04-0519.33%
2026-09-01
Source-checked
35Llama 4 ScoutMeta2026-04-0514%
2026-09-01
Source-checked
36Command ACohere2026-03-1513%
2026-09-01
Source-checked
37Claude 3 OpusAnthropic2024-03-043.33%
2026-09-01
Source-checked
Method

What it covers

Archived — GPT-5.2 and peers now report 100% on recent AIME papers with extended thinking. Treat as solved, not a frontier separator. Prefer /benchmarks/frontiermath.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Tools that report it

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does AIME measure?

Hard high-school competition math: 15 integer-answer problems per exam. Adopted as a frontier reasoning eval because each year's fresh paper resists contamination — until it leaks.

What does a high AIME score not prove?

A strong AIME result says nothing about:

  • Tiny sample (15 problems/year) — one problem is ~6.7%, so single-run scores are noisy; demand pass@k or multiple seeds.
  • General capability — strong AIME scores correlate with reasoning training, not with everyday helpfulness.
  • Which year — '90% on AIME' is meaningless without the exam year and number of attempts.

How is AIME scored?

Accuracy (often pass@1 or cons@64). Task format: 15 problems, integer answers 0-999; reported as accuracy on AIME 2024/2025 papers.

Is AIME saturated?

AIME is currently marked Defunct in the atlas.

Can AIME results be contaminated by training data?

Contamination risk for AIME is graded Medium contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.