Why this benchmark is useful
Historical competition-math suite — once a standard reasoning yardstick. Now saturated; keep for context, prefer harder math evals for frontier separation.
Step-by-step solutions to 12,500 competition mathematics problems (AMC/AIME-style), graded on the final boxed answer.
Historical competition-math suite — once a standard reasoning yardstick. Now saturated; keep for context, prefer harder math evals for frontier separation.
Final-answer accuracy only — wrong work with a right box still scores. MATH-500 and full MATH cluster near ceiling for frontier models; not for current ranking.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | GPT-5 (high)OpenAI | 2025-08-07 | 99.4% | 2026-09-01 | Source-checked |
| 2 | OpenAI o3OpenAI | 2025-04-16 | 99.2% | 2026-09-01 | Source-checked |
| 3 | GPT-5 (medium)OpenAI | 2025-08-07 | 99.13% | 2026-09-01 | Source-checked |
| 4 | Grok 4xAI | 2025-07-10 | 99% | 2026-09-01 | Source-checked |
| 5 | GPT-5 (low)OpenAI | 2025-08-07 | 98.73% | 2026-09-01 | Source-checked |
| 6 | Claude 4 Opus (Reasoning)Anthropic | 2025-05-22 | 98.2% | 2026-09-01 | Source-checked |
| 7 | OpenAI o1OpenAI | 2024-09-12 | 97% | 2026-09-01 | Source-checked |
| 1 | o3 (high compute)OpenAIMATH (Hendrycks et al.). Extended thinking. | 2025-04-16 | 91.1% | OpenAI: Learning to Reason2025-01-20 | Source-checked |
| 2 | GPT 5.5OpenAIMATH. Contamination warning applies. | 2026-04-23 | 89.6% | OpenAI model release notes2026-04-23 | Needs audit |
| 10 | Llama 4 MaverickMeta | 2026-04-05 | 88.87% | 2026-09-01 | Source-checked |
| 3 | Claude Opus 4.8AnthropicMATH (full set). Contamination warning applies. | 2026-05-28 | 88.4% | Anthropic Claude 4 family2026-05-28 | Needs audit |
| 4 | GPT 5OpenAIMATH. Contamination warning applies. | 2025-12-11 | 87.2% | OpenAI model release notes2025-12-01 | Needs audit |
| 5 | Gemini 3.1 ProGoogleMATH. Contamination warning applies. | 2026-02-19 | 86.7% | Google Gemini 3.1 Pro2026-02-19 | Needs audit |
| 6 | Claude Opus 4AnthropicMATH (full 12,500 problems). High contamination risk — these numbers reflect trained performance, not raw reasoning. | 2024-07-01 | 85.6% | Anthropic model card2025-06-01 | Needs audit |
| 15 | Llama 4 ScoutMeta | 2026-04-05 | 84.4% | 2026-09-01 | Source-checked |
| 7 | Claude Sonnet 4.6AnthropicMATH. Contamination warning applies. | 2026-02-17 | 84.2% | Anthropic Claude 3.7/4.6 model card2026-02-17 | Needs audit |
| 8 | DeepSeek V4 Pro MaxDeepSeekMATH. Contamination warning applies. | 2026-04-24 | 83.9% | DeepSeek V4 release2026-04-24 | Needs audit |
| 9 | Kimi K2.6MoonshotMATH. Contamination warning applies. | 2026-04-20 | 82.5% | Moonshot Kimi K2.62026-04-20 | Needs audit |
| 19 | Command ACohere | 2026-03-15 | 81.87% | 2026-09-01 | Source-checked |
| 10 | Qwen 3.7 MaxAlibabaMATH. Contamination warning applies. | 2026-05-20 | 81.6% | Alibaba Qwen 3.7 Max2026-05-20 | Needs audit |
| 11 | DeepSeek R1DeepSeekMATH. Contamination warning applies. | 2025-01-20 | 79.8% | DeepSeek R1 model card2025-01-20 | Needs audit |
| 22 | DeepSeek Coder V2DeepSeek | 2024-05-15 | 74.27% | 2026-09-01 | Source-checked |
| 12 | Claude 3.5 Sonnet (2024-06)Anthropic | 2024-06-20 | 71.1% | Anthropic Claude 3.5 Sonnet2024-06-20 | Source-checked |
| 24 | Claude 3 OpusAnthropic | 2024-03-04 | 64.07% | 2026-09-01 | Source-checked |
| 13 | GPT-4 (2023)OpenAI | 2023-03-14 | 52.9% | GPT-4 Technical Report2023-03-15 | Source-checked |
Archived — MATH-500 (AA public run) now reports 99%+ for frontier models. Full MATH set clusters above 90%. Prefer /benchmarks/frontiermath for hard math separation.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
Step-by-step solutions to 12,500 competition mathematics problems (AMC/AIME-style), graded on the final boxed answer.
A strong MATH result says nothing about:
Accuracy on the final answer. Task format: 12,500 problems across 7 subjects and 5 difficulty levels; answer-matching on a normalized final result.
MATH is currently marked Defunct in the atlas.
Contamination risk for MATH is graded High contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.