Benchmarks / Math

MATH

Step-by-step solutions to 12,500 competition mathematics problems (AMC/AIME-style), graded on the final boxed answer.

MathSolidDefunctHigh contamination riskSince 2021
What this does not measure
  • Whether the reasoning is correct — only the final answer is graded, so a right answer from wrong work still scores.
  • Novel mathematics — every problem has a known closed-form answer from existing competitions.
Analysis

Why this benchmark is useful

It gives a crisp reasoning signal with right-or-wrong answers. Use it to compare structured problem solving, while watching sample size and pass@k settings.

Scope

Coverage map

Task family
Math
Format
12,500 problems across 7 subjects and 5 difficulty levels; answer-matching on a normalized final result.
Scoring
Accuracy on the final answer.
Maintainer
Hendrycks et al.
Reading guide

How to read the scores

Read MATH as a defunct signal with high contamination risk. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.

Blind spots

What it does not cover

  • Whether the reasoning is correct — only the final answer is graded, so a right answer from wrong work still scores.
  • Novel mathematics — every problem has a known closed-form answer from existing competitions.
Scores

Evidence ledger

20 rows
2020 rows
99.2%best score
10source-checked
3sources
2026-07-21to 2023-03-15
7

api · api

Papers / model cards4

manual-snapshot · manual

Vendor claims9

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. OpenAI o399.2% ·
  2. OpenAI o197% ·
  3. Llama 4 Maverick88.87% ·
  4. Llama 4 Scout84.4% ·
  5. Command A81.87% ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1OpenAI o3OpenAI2025-04-1699.2%
2026-07-21
Source-checked
2OpenAI o1OpenAI2024-09-1297%
2026-07-21
Source-checked
1o3 (high compute)OpenAIMATH (Hendrycks et al.). Extended thinking.2025-04-1691.1%
OpenAI: Learning to Reason2025-01-20
Source-checked
2GPT 5.5OpenAIMATH. Contamination warning applies.2026-04-2389.6%
OpenAI model release notes2026-04-23
Needs audit
5Llama 4 MaverickMeta2026-04-0588.87%
2026-07-21
Source-checked
3Claude Opus 4.8AnthropicMATH (full set). Contamination warning applies.2026-05-2888.4%
Anthropic Claude 4 family2026-05-28
Needs audit
4GPT 5OpenAIMATH. Contamination warning applies.2025-12-1187.2%
OpenAI model release notes2025-12-01
Needs audit
5Gemini 3.1 ProGoogleMATH. Contamination warning applies.2026-02-1986.7%
Google Gemini 3.1 Pro2026-02-19
Needs audit
6Claude Opus 4AnthropicMATH (full 12,500 problems). High contamination risk — these numbers reflect trained performance, not raw reasoning.2024-07-0185.6%
Anthropic model card2025-06-01
Needs audit
10Llama 4 ScoutMeta2026-04-0584.4%
2026-07-21
Source-checked
7Claude Sonnet 4.6AnthropicMATH. Contamination warning applies.2026-02-1784.2%
Anthropic Claude 3.7/4.6 model card2026-02-17
Needs audit
8DeepSeek V4 Pro MaxDeepSeekMATH. Contamination warning applies.2026-04-2483.9%
DeepSeek V4 release2026-04-24
Needs audit
9Kimi K2.6MoonshotMATH. Contamination warning applies.2026-04-2082.5%
Moonshot Kimi K2.62026-04-20
Needs audit
14Command ACohere2026-03-1581.87%
2026-07-21
Source-checked
10Qwen 3.7 MaxAlibabaMATH. Contamination warning applies.2026-05-2081.6%
Alibaba Qwen 3.7 Max2026-05-20
Needs audit
11DeepSeek R1DeepSeekMATH. Contamination warning applies.2025-01-2079.8%
DeepSeek R1 model card2025-01-20
Needs audit
17DeepSeek Coder V2DeepSeek2024-05-1574.27%
2026-07-21
Source-checked
12Claude 3.5 Sonnet (2024-06)Anthropic2024-06-2071.1%
Anthropic Claude 3.5 Sonnet2024-06-20
Source-checked
19Claude 3 OpusAnthropic2024-03-0464.07%
2026-07-21
Source-checked
13GPT-4 (2023)OpenAI2023-03-1452.9%
GPT-4 Technical Report2023-03-15
Source-checked
Method

What it covers

Archived — MATH-500 (AA public run) now reports 99%+ for frontier models. Full MATH set clusters above 90%. Prefer FrontierMath for hard math separation.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Tools that report it

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.