Why this benchmark is useful
It gives a crisp reasoning signal with right-or-wrong answers. Use it to compare structured problem solving, while watching sample size and pass@k settings.
Benchmarks / Math
Step-by-step solutions to 12,500 competition mathematics problems (AMC/AIME-style), graded on the final boxed answer.
It gives a crisp reasoning signal with right-or-wrong answers. Use it to compare structured problem solving, while watching sample size and pass@k settings.
Read MATH as a defunct signal with high contamination risk. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | OpenAI o3OpenAI | 2025-04-16 | 99.2% | 2026-07-21 | Source-checked |
| 2 | OpenAI o1OpenAI | 2024-09-12 | 97% | 2026-07-21 | Source-checked |
| 1 | o3 (high compute)OpenAIMATH (Hendrycks et al.). Extended thinking. | 2025-04-16 | 91.1% | OpenAI: Learning to Reason2025-01-20 | Source-checked |
| 2 | GPT 5.5OpenAIMATH. Contamination warning applies. | 2026-04-23 | 89.6% | OpenAI model release notes2026-04-23 | Needs audit |
| 5 | Llama 4 MaverickMeta | 2026-04-05 | 88.87% | 2026-07-21 | Source-checked |
| 3 | Claude Opus 4.8AnthropicMATH (full set). Contamination warning applies. | 2026-05-28 | 88.4% | Anthropic Claude 4 family2026-05-28 | Needs audit |
| 4 | GPT 5OpenAIMATH. Contamination warning applies. | 2025-12-11 | 87.2% | OpenAI model release notes2025-12-01 | Needs audit |
| 5 | Gemini 3.1 ProGoogleMATH. Contamination warning applies. | 2026-02-19 | 86.7% | Google Gemini 3.1 Pro2026-02-19 | Needs audit |
| 6 | Claude Opus 4AnthropicMATH (full 12,500 problems). High contamination risk — these numbers reflect trained performance, not raw reasoning. | 2024-07-01 | 85.6% | Anthropic model card2025-06-01 | Needs audit |
| 10 | Llama 4 ScoutMeta | 2026-04-05 | 84.4% | 2026-07-21 | Source-checked |
| 7 | Claude Sonnet 4.6AnthropicMATH. Contamination warning applies. | 2026-02-17 | 84.2% | Anthropic Claude 3.7/4.6 model card2026-02-17 | Needs audit |
| 8 | DeepSeek V4 Pro MaxDeepSeekMATH. Contamination warning applies. | 2026-04-24 | 83.9% | DeepSeek V4 release2026-04-24 | Needs audit |
| 9 | Kimi K2.6MoonshotMATH. Contamination warning applies. | 2026-04-20 | 82.5% | Moonshot Kimi K2.62026-04-20 | Needs audit |
| 14 | Command ACohere | 2026-03-15 | 81.87% | 2026-07-21 | Source-checked |
| 10 | Qwen 3.7 MaxAlibabaMATH. Contamination warning applies. | 2026-05-20 | 81.6% | Alibaba Qwen 3.7 Max2026-05-20 | Needs audit |
| 11 | DeepSeek R1DeepSeekMATH. Contamination warning applies. | 2025-01-20 | 79.8% | DeepSeek R1 model card2025-01-20 | Needs audit |
| 17 | DeepSeek Coder V2DeepSeek | 2024-05-15 | 74.27% | 2026-07-21 | Source-checked |
| 12 | Claude 3.5 Sonnet (2024-06)Anthropic | 2024-06-20 | 71.1% | Anthropic Claude 3.5 Sonnet2024-06-20 | Source-checked |
| 19 | Claude 3 OpusAnthropic | 2024-03-04 | 64.07% | 2026-07-21 | Source-checked |
| 13 | GPT-4 (2023)OpenAI | 2023-03-14 | 52.9% | GPT-4 Technical Report2023-03-15 | Source-checked |
Archived — MATH-500 (AA public run) now reports 99%+ for frontier models. Full MATH set clusters above 90%. Prefer FrontierMath for hard math separation.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
The Pack · Editorial newsletter
One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.