Benchmarks / Math

MATH

Step-by-step solutions to 12,500 competition mathematics problems (AMC/AIME-style), graded on the final boxed answer.

What this does not measure
  • Whether the reasoning is correct — only the final answer is graded, so a right answer from wrong work still scores.
  • Novel mathematics — every problem has a known closed-form answer from existing competitions.
Analysis

Why this benchmark is useful

Historical competition-math suite — once a standard reasoning yardstick. Now saturated; keep for context, prefer harder math evals for frontier separation.

Scope

Coverage map

Task family
Math
Format
12,500 problems across 7 subjects and 5 difficulty levels; answer-matching on a normalized final result.
Scoring
Accuracy on the final answer.
Maintainer
Hendrycks et al.
Reading guide

How to read the scores

Final-answer accuracy only — wrong work with a right box still scores. MATH-500 and full MATH cluster near ceiling for frontier models; not for current ranking.

Blind spots

What it does not cover

  • Whether the reasoning is correct — only the final answer is graded, so a right answer from wrong work still scores.
  • Novel mathematics — every problem has a known closed-form answer from existing competitions.
Scores

Evidence ledger

25 rows
2525 rows
99.4%best score
15source-checked
3sources
2026-09-01to 2023-03-15
12

api · api

Papers / model cards4

manual-snapshot · manual

Vendor claims9

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. GPT-5 (high)99.4% ·
  2. OpenAI o399.2% ·
  3. GPT-5 (medium)99.13% ·
  4. Grok 499% ·
  5. GPT-5 (low)98.73% ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1GPT-5 (high)OpenAI2025-08-0799.4%
2026-09-01
Source-checked
2OpenAI o3OpenAI2025-04-1699.2%
2026-09-01
Source-checked
3GPT-5 (medium)OpenAI2025-08-0799.13%
2026-09-01
Source-checked
4Grok 4xAI2025-07-1099%
2026-09-01
Source-checked
5GPT-5 (low)OpenAI2025-08-0798.73%
2026-09-01
Source-checked
6Claude 4 Opus (Reasoning)Anthropic2025-05-2298.2%
2026-09-01
Source-checked
7OpenAI o1OpenAI2024-09-1297%
2026-09-01
Source-checked
1o3 (high compute)OpenAIMATH (Hendrycks et al.). Extended thinking.2025-04-1691.1%
OpenAI: Learning to Reason2025-01-20
Source-checked
2GPT 5.5OpenAIMATH. Contamination warning applies.2026-04-2389.6%
OpenAI model release notes2026-04-23
Needs audit
10Llama 4 MaverickMeta2026-04-0588.87%
2026-09-01
Source-checked
3Claude Opus 4.8AnthropicMATH (full set). Contamination warning applies.2026-05-2888.4%
Anthropic Claude 4 family2026-05-28
Needs audit
4GPT 5OpenAIMATH. Contamination warning applies.2025-12-1187.2%
OpenAI model release notes2025-12-01
Needs audit
5Gemini 3.1 ProGoogleMATH. Contamination warning applies.2026-02-1986.7%
Google Gemini 3.1 Pro2026-02-19
Needs audit
6Claude Opus 4AnthropicMATH (full 12,500 problems). High contamination risk — these numbers reflect trained performance, not raw reasoning.2024-07-0185.6%
Anthropic model card2025-06-01
Needs audit
15Llama 4 ScoutMeta2026-04-0584.4%
2026-09-01
Source-checked
7Claude Sonnet 4.6AnthropicMATH. Contamination warning applies.2026-02-1784.2%
Anthropic Claude 3.7/4.6 model card2026-02-17
Needs audit
8DeepSeek V4 Pro MaxDeepSeekMATH. Contamination warning applies.2026-04-2483.9%
DeepSeek V4 release2026-04-24
Needs audit
9Kimi K2.6MoonshotMATH. Contamination warning applies.2026-04-2082.5%
Moonshot Kimi K2.62026-04-20
Needs audit
19Command ACohere2026-03-1581.87%
2026-09-01
Source-checked
10Qwen 3.7 MaxAlibabaMATH. Contamination warning applies.2026-05-2081.6%
Alibaba Qwen 3.7 Max2026-05-20
Needs audit
11DeepSeek R1DeepSeekMATH. Contamination warning applies.2025-01-2079.8%
DeepSeek R1 model card2025-01-20
Needs audit
22DeepSeek Coder V2DeepSeek2024-05-1574.27%
2026-09-01
Source-checked
12Claude 3.5 Sonnet (2024-06)Anthropic2024-06-2071.1%
Anthropic Claude 3.5 Sonnet2024-06-20
Source-checked
24Claude 3 OpusAnthropic2024-03-0464.07%
2026-09-01
Source-checked
13GPT-4 (2023)OpenAI2023-03-1452.9%
GPT-4 Technical Report2023-03-15
Source-checked
Method

What it covers

Archived — MATH-500 (AA public run) now reports 99%+ for frontier models. Full MATH set clusters above 90%. Prefer /benchmarks/frontiermath for hard math separation.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Tools that report it

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does MATH measure?

Step-by-step solutions to 12,500 competition mathematics problems (AMC/AIME-style), graded on the final boxed answer.

What does a high MATH score not prove?

A strong MATH result says nothing about:

  • Whether the reasoning is correct — only the final answer is graded, so a right answer from wrong work still scores.
  • Novel mathematics — every problem has a known closed-form answer from existing competitions.

How is MATH scored?

Accuracy on the final answer. Task format: 12,500 problems across 7 subjects and 5 difficulty levels; answer-matching on a normalized final result.

Is MATH saturated?

MATH is currently marked Defunct in the atlas.

Can MATH results be contaminated by training data?

Contamination risk for MATH is graded High contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.