Benchmarks / Math

FrontierMath

FrontierMath Tiers 1-3

Hard unpublished mathematics: 295 private Tiers 1-3 problems (v2) where the model writes a Python answer() after iterative reasoning and code execution.

What this does not measure
  • Tier 4 research-level problems — that 43-problem expansion is a separate, harder set. Do not mix Tiers 1-3 headlines with Tier 4.
  • Contest MATH / AIME skill — those suites are saturated. FrontierMath is original expert-authored problems, not recycled olympiad papers.
  • Whether the derivation is human-readable — scoring is pass/fail on the submitted Python object, not a proof grade.
Analysis

Why this benchmark is useful

When MATH and AIME no longer separate the frontier, this is the live math discriminator: original expert problems, Python-enabled, still far from ceiling.

Scope

Coverage map

Task family
Math
Format
Private set of 295 Tiers 1-3 problems (v2, 12 Jun 2026 correction). The model may think aloud, call a stateless Python tool, and must submit an answer() function. Token cap 1,000,000.
Scoring
Pass rate on the private Tiers 1-3 v2 set (1 if the submitted object matches, else 0). VerdictPal rows use Epoch's best internal run per model from the public CSV.
Maintainer
Epoch AI
Reading guide

How to read the scores

Read the private Tiers 1-3 v2 pass rate, not vendor MATH-500 headlines. Epoch funds include OpenAI exclusive access to a subset — treat OpenAI rows with that conflict in mind. GPT-5.6 Sol is not in the 2026-08-15 Epoch CSV yet.

Blind spots

What it does not cover

  • Tier 4 research-level problems — that 43-problem expansion is a separate, harder set. Do not mix Tiers 1-3 headlines with Tier 4.
  • Contest MATH / AIME skill — those suites are saturated. FrontierMath is original expert-authored problems, not recycled olympiad papers.
  • Whether the derivation is human-readable — scoring is pass/fail on the submitted Python object, not a proof grade.
Scores

Evidence ledger

295 private problems after the 12 Jun 2026 correction. This is the hub primary.

15 rows
1515 rows
52.4%best score
15source-checked
1sources
2026-08-15source date
Papers / model cards15

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. GPT-5.5 Pro52.4% · Epoch AI FrontierMath Tiers 1-3 (v2) CSV
  2. GPT 5.551.7% · Epoch AI FrontierMath Tiers 1-3 (v2) CSV
  3. GPT-5.4 Pro50% · Epoch AI FrontierMath Tiers 1-3 (v2) CSV
  4. GPT 5.447.6% · Epoch AI FrontierMath Tiers 1-3 (v2) CSV
  5. Claude Opus 4.847.24% · Epoch AI FrontierMath Tiers 1-3 (v2) CSV

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1GPT-5.5 ProOpenAIEpoch CSV row `gpt-5.5-pro-pre-release_high`; best internal Tiers 1-3 v2 run (started 2026-04-23). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note.2026-04-2352.4%
Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15
Source-checked
2GPT 5.5OpenAIEpoch CSV row `gpt-5.5-pre-release_xhigh`; best internal Tiers 1-3 v2 run (started 2026-04-23). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note.2026-04-2351.7%
Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15
Source-checked
3GPT-5.4 ProOpenAIEpoch CSV row `gpt-5.4-pro-2026-03-05_xhigh`; best internal Tiers 1-3 v2 run (started 2026-03-06). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note.2026-03-0550%
Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15
Source-checked
4GPT 5.4OpenAIEpoch CSV row `gpt-5.4-2026-03-05_xhigh`; best internal Tiers 1-3 v2 run (started 2026-03-06). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note.2026-03-0547.6%
Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15
Source-checked
5Claude Opus 4.8AnthropicEpoch CSV row `claude-opus-4-8_max`; best internal Tiers 1-3 v2 run (started 2026-06-08). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note.2026-05-2847.24%
Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15
Source-checked
6Claude Opus 4.7AnthropicEpoch CSV row `claude-opus-4-7_xhigh`; best internal Tiers 1-3 v2 run (started 2026-04-17). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note.2026-04-1643.79%
Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15
Source-checked
7Claude Opus 4.6AnthropicEpoch CSV row `claude-opus-4-6_max`; best internal Tiers 1-3 v2 run (started 2026-02-12). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note.2026-02-0540.7%
Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15
Source-checked
8GPT-5.2OpenAIEpoch CSV row `gpt-5.2-2025-12-11_xhigh`; best internal Tiers 1-3 v2 run (started 2025-12-13). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note.2025-12-1140.7%
Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15
Source-checked
9Muse SparkOtherEpoch CSV row `muse-spark`; best internal Tiers 1-3 v2 run (started 2026-04-08). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note.2026-01-0139%
Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15
Source-checked
10Gemini 3.5 FlashGoogleEpoch CSV row `gemini-3.5-flash_high`; best internal Tiers 1-3 v2 run (started 2026-05-22). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note.2026-05-1938.97%
Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15
Source-checked
11Kimi K2.6MoonshotEpoch CSV row `kimi-k2.6`; best internal Tiers 1-3 v2 run (started 2026-05-07). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note.2026-04-2038.97%
Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15
Source-checked
12Gemini 3 ProGoogleEpoch CSV row `gemini-3-pro-preview`; best internal Tiers 1-3 v2 run (started 2025-11-21). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note.2025-11-1837.6%
Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15
Source-checked
13Gemini 3.1 ProGoogleEpoch CSV row `gemini-3.1-pro-preview`; best internal Tiers 1-3 v2 run (started 2026-02-19). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note.2026-02-1936.9%
Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15
Source-checked
14GLM-5.1ZhipuEpoch CSV row `glm-5.1`; best internal Tiers 1-3 v2 run (started 2026-05-11). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note.2026-04-0733.45%
Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15
Source-checked
15Claude Sonnet 4.6AnthropicEpoch CSV row `claude-sonnet-4-6_16K`; best internal Tiers 1-3 v2 run (started 2026-02-20). OpenAI funded FrontierMath. Epoch publishes a conflict-of-interest note.2026-02-1732.4%
Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15
Source-checked
Method

What it covers

Scores ingested from Epoch's 2026-08-15 public CSV (internal runs). One row per atlas model, best effort variant. OpenAI funded FrontierMath; Epoch publishes a conflict-of-interest note.

Score ceiling

Where it breaks down

Human expert time is hours to days per problem. Frontier pass rates in the Epoch CSV cluster well below 60% on Tiers 1-3 v2.

Tools that report it

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does FrontierMath measure?

Hard unpublished mathematics: 295 private Tiers 1-3 problems (v2) where the model writes a Python answer() after iterative reasoning and code execution.

What does a high FrontierMath score not prove?

A strong FrontierMath result says nothing about:

  • Tier 4 research-level problems — that 43-problem expansion is a separate, harder set. Do not mix Tiers 1-3 headlines with Tier 4.
  • Contest MATH / AIME skill — those suites are saturated. FrontierMath is original expert-authored problems, not recycled olympiad papers.
  • Whether the derivation is human-readable — scoring is pass/fail on the submitted Python object, not a proof grade.

How is FrontierMath scored?

Pass rate on the private Tiers 1-3 v2 set (1 if the submitted object matches, else 0). VerdictPal rows use Epoch's best internal run per model from the public CSV. Task format: Private set of 295 Tiers 1-3 problems (v2, 12 Jun 2026 correction). The model may think aloud, call a stateless Python tool, and must submit an answer() function. Token cap 1,000,000.

Is FrontierMath saturated?

FrontierMath is currently marked Active in the atlas. Ceiling context: Human expert time is hours to days per problem. Frontier pass rates in the Epoch CSV cluster well below 60% on Tiers 1-3 v2.

Can FrontierMath results be contaminated by training data?

Contamination risk for FrontierMath is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.