Why this benchmark is useful
When MATH and AIME no longer separate the frontier, this is the live math discriminator: original expert problems, Python-enabled, still far from ceiling.
FrontierMath Tiers 1-3
Hard unpublished mathematics: 295 private Tiers 1-3 problems (v2) where the model writes a Python answer() after iterative reasoning and code execution.
When MATH and AIME no longer separate the frontier, this is the live math discriminator: original expert problems, Python-enabled, still far from ceiling.
Read the private Tiers 1-3 v2 pass rate, not vendor MATH-500 headlines. Epoch funds include OpenAI exclusive access to a subset — treat OpenAI rows with that conflict in mind. GPT-5.6 Sol is not in the 2026-08-15 Epoch CSV yet.
295 private problems after the 12 Jun 2026 correction. This is the hub primary.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | GPT-5.5 ProOpenAIEpoch CSV row `gpt-5.5-pro-pre-release_high`; best internal Tiers 1-3 v2 run (started 2026-04-23). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note. | 2026-04-23 | 52.4% | Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15 | Source-checked |
| 2 | GPT 5.5OpenAIEpoch CSV row `gpt-5.5-pre-release_xhigh`; best internal Tiers 1-3 v2 run (started 2026-04-23). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note. | 2026-04-23 | 51.7% | Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15 | Source-checked |
| 3 | GPT-5.4 ProOpenAIEpoch CSV row `gpt-5.4-pro-2026-03-05_xhigh`; best internal Tiers 1-3 v2 run (started 2026-03-06). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note. | 2026-03-05 | 50% | Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15 | Source-checked |
| 4 | GPT 5.4OpenAIEpoch CSV row `gpt-5.4-2026-03-05_xhigh`; best internal Tiers 1-3 v2 run (started 2026-03-06). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note. | 2026-03-05 | 47.6% | Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15 | Source-checked |
| 5 | Claude Opus 4.8AnthropicEpoch CSV row `claude-opus-4-8_max`; best internal Tiers 1-3 v2 run (started 2026-06-08). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note. | 2026-05-28 | 47.24% | Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15 | Source-checked |
| 6 | Claude Opus 4.7AnthropicEpoch CSV row `claude-opus-4-7_xhigh`; best internal Tiers 1-3 v2 run (started 2026-04-17). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note. | 2026-04-16 | 43.79% | Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15 | Source-checked |
| 7 | Claude Opus 4.6AnthropicEpoch CSV row `claude-opus-4-6_max`; best internal Tiers 1-3 v2 run (started 2026-02-12). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note. | 2026-02-05 | 40.7% | Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15 | Source-checked |
| 8 | GPT-5.2OpenAIEpoch CSV row `gpt-5.2-2025-12-11_xhigh`; best internal Tiers 1-3 v2 run (started 2025-12-13). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note. | 2025-12-11 | 40.7% | Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15 | Source-checked |
| 9 | Muse SparkOtherEpoch CSV row `muse-spark`; best internal Tiers 1-3 v2 run (started 2026-04-08). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note. | 2026-01-01 | 39% | Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15 | Source-checked |
| 10 | Gemini 3.5 FlashGoogleEpoch CSV row `gemini-3.5-flash_high`; best internal Tiers 1-3 v2 run (started 2026-05-22). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note. | 2026-05-19 | 38.97% | Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15 | Source-checked |
| 11 | Kimi K2.6MoonshotEpoch CSV row `kimi-k2.6`; best internal Tiers 1-3 v2 run (started 2026-05-07). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note. | 2026-04-20 | 38.97% | Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15 | Source-checked |
| 12 | Gemini 3 ProGoogleEpoch CSV row `gemini-3-pro-preview`; best internal Tiers 1-3 v2 run (started 2025-11-21). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note. | 2025-11-18 | 37.6% | Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15 | Source-checked |
| 13 | Gemini 3.1 ProGoogleEpoch CSV row `gemini-3.1-pro-preview`; best internal Tiers 1-3 v2 run (started 2026-02-19). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note. | 2026-02-19 | 36.9% | Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15 | Source-checked |
| 14 | GLM-5.1ZhipuEpoch CSV row `glm-5.1`; best internal Tiers 1-3 v2 run (started 2026-05-11). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note. | 2026-04-07 | 33.45% | Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15 | Source-checked |
| 15 | Claude Sonnet 4.6AnthropicEpoch CSV row `claude-sonnet-4-6_16K`; best internal Tiers 1-3 v2 run (started 2026-02-20). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note. | 2026-02-17 | 32.4% | Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15 | Source-checked |
Scores ingested from Epoch's 2026-08-15 public CSV (internal runs). One row per atlas model, best effort variant. OpenAI funded FrontierMath; Epoch publishes a conflict-of-interest note.
Human expert time is hours to days per problem. Frontier pass rates in the Epoch CSV cluster well below 60% on Tiers 1-3 v2.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
Hard unpublished mathematics: 295 private Tiers 1-3 problems (v2) where the model writes a Python answer() after iterative reasoning and code execution.
A strong FrontierMath result says nothing about:
Pass rate on the private Tiers 1-3 v2 set (1 if the submitted object matches, else 0). VerdictPal rows use Epoch's best internal run per model from the public CSV. Task format: Private set of 295 Tiers 1-3 problems (v2, 12 Jun 2026 correction). The model may think aloud, call a stateless Python tool, and must submit an answer() function. Token cap 1,000,000.
FrontierMath is currently marked Active in the atlas. Ceiling context: Human expert time is hours to days per problem. Frontier pass rates in the Epoch CSV cluster well below 60% on Tiers 1-3 v2.
Contamination risk for FrontierMath is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.