Benchmarks / Math

American Invitational Mathematics Examination

AIME

Hard high-school competition math: 15 integer-answer problems per exam. Adopted as a frontier reasoning eval because each year's fresh paper resists contamination — until it leaks.

MathSolidDefunctMedium contamination riskSince 2024
What this does not measure
  • Tiny sample (15 problems/year) — one problem is ~6.7%, so single-run scores are noisy; demand pass@k or multiple seeds.
  • General capability — strong AIME scores correlate with reasoning training, not with everyday helpfulness.
  • Which year — '90% on AIME' is meaningless without the exam year and number of attempts.
Analysis

Why this benchmark is useful

It gives a crisp reasoning signal with right-or-wrong answers. Use it to compare structured problem solving, while watching sample size and pass@k settings.

Scope

Coverage map

Task family
Math
Format
15 problems, integer answers 0-999; reported as accuracy on AIME 2024/2025 papers.
Scoring
Accuracy (often pass@1 or cons@64).
Maintainer
MAA (exam) · used as an LLM eval
Reading guide

How to read the scores

Read AIME as a defunct signal with medium contamination risk. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.

Blind spots

What it does not cover

  • Tiny sample (15 problems/year) — one problem is ~6.7%, so single-run scores are noisy; demand pass@k or multiple seeds.
  • General capability — strong AIME scores correlate with reasoning training, not with everyday helpfulness.
  • Which year — '90% on AIME' is meaningless without the exam year and number of attempts.
Scores

Evidence ledger

19 rows
1919 rows
95.67%best score
13source-checked
3sources
2026-07-21to 2024-09-12
10

api · api

Papers / model cards2

manual-snapshot · manual

Vendor claims7

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Gemini 3 Pro95.67% ·
  2. GPT-5.194% ·
  3. OpenAI o388.33% ·
  4. Claude Opus 4.562.67% ·
  5. Gemini 3 Flash Preview55.67% ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Gemini 3 ProGoogle2025-11-1895.67%
2026-07-21
Source-checked
1Claude Opus 4.8AnthropicAIME 2024 with extended thinking.2026-05-2894.2%
Anthropic Claude 4 family2026-05-28
Needs audit
3GPT-5.1OpenAI2025-11-1394%
2026-07-21
Source-checked
2Claude 3.7 Sonnet (2025-02)AnthropicAIME 2024 with extended thinking enabled.2025-02-1993.7%
Anthropic Claude 3.7 Sonnet announcement2025-02-19
Source-checked
3GPT 5.5OpenAI2026-04-2391.8%
OpenAI model release notes2026-04-23
Needs audit
4GPT 5OpenAI2025-12-1189.2%
OpenAI model release notes2025-12-01
Needs audit
5DeepSeek R1DeepSeekAIME 2024 with chain-of-thought.2025-01-2088.4%
DeepSeek R1 model card2025-01-20
Needs audit
8OpenAI o3OpenAI2025-04-1688.33%
2026-07-21
Source-checked
6Gemini 3.1 ProGoogle2026-02-1987.6%
Google Gemini 3.1 Pro2026-02-19
Needs audit
7o3 (cons@64, AIME 2025)OpenAIAIME 2025 paper; 64-sample consensus. Single-shot is markedly lower.2025-04-1687.5%
OpenAI: Learning to Reason2025-01-20
Source-checked
9Kimi K2.6Moonshot2026-04-2085.2%
Moonshot Kimi K2.62026-04-20
Needs audit
10OpenAI o1 (AIME 2024, cons@64)OpenAIWith 64-sample consensus; single-shot is markedly lower.2024-09-1283.3%
OpenAI: Learning to Reason2024-09-12
Source-checked
13Claude Opus 4.5Anthropic2025-11-2462.67%
2026-07-21
Source-checked
14Gemini 3 Flash PreviewGoogle2025-12-1755.67%
2026-07-21
Source-checked
15Mistral Large 3Mistral2025-12-0238%
2026-07-21
Source-checked
16Llama 4 MaverickMeta2026-04-0519.33%
2026-07-21
Source-checked
17Llama 4 ScoutMeta2026-04-0514%
2026-07-21
Source-checked
18Command ACohere2026-03-1513%
2026-07-21
Source-checked
19Claude 3 OpusAnthropic2024-03-043.33%
2026-07-21
Source-checked
Method

What it covers

Archived — GPT-5.2 and peers now report 100% on recent AIME papers with extended thinking. Treat as solved, not a frontier separator.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Tools that report it

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.