Benchmarks / Reasoning

CritPt

Critical Physics Tasks

Unpublished research-level physics: 71 composite challenges (70 test + 1 example) written by 50+ active physicists across 11 subfields, plus 190 simpler checkpoint tasks. 6% weight in AA Intelligence Index v4.1.

What this does not measure
  • Textbook or contest physics — problems are original research-scale tasks, not recycled qualifying-exam items.
  • Whether the derivation is a readable paper — scoring is machine-verifiable on guess-resistant formats (arrays, symbolic expressions, Python functions), not a proof grade.
  • Cross-domain scientific reasoning — CritPt is physics only. Use HLE or SciCode for broader science.
Analysis

Why this benchmark is useful

When GPQA and HLE no longer tell you whether a model can do research-scale physics: original, unpublished problems with guess-resistant answers, still far from ceiling (frontier around 30%).

Scope

Coverage map

Task family
Reasoning
Format
71 composite research challenges spanning condensed matter, quantum, AMO, astrophysics, HEP, mathematical physics, statistical physics, nuclear, nonlinear dynamics, fluids, and biophysics. Models may use coding tools in some AA runs; checkpoint tasks (190) are a separate, easier slice.
Scoring
Average accuracy on the 70-test composite set. Automated grading for physics-specific output formats. VerdictPal rows use Artificial Analysis' independent CritPt run.
Maintainer
Argonne / UIUC / Artificial Analysis
Reading guide

How to read the scores

Read the composite challenge accuracy, not checkpoint-only headlines. GPT-5.6 Sol (max) leads the 1 Sep 2026 AA board at 32.3%. Tool-augmented runs are a separate, higher tier — do not mix them with base rows.

Blind spots

What it does not cover

  • Textbook or contest physics — problems are original research-scale tasks, not recycled qualifying-exam items.
  • Whether the derivation is a readable paper — scoring is machine-verifiable on guess-resistant formats (arrays, symbolic expressions, Python functions), not a proof grade.
  • Cross-domain scientific reasoning — CritPt is physics only. Use HLE or SciCode for broader science.
Scores

Evidence ledger

12 rows
1212 rows
32.3%best score
12source-checked
1sources
2026-09-01source date
12

api · api

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. GPT-5.6 Sol32.3% ·
  2. GPT-5.5 Pro30.6% ·
  3. GPT-5.6 Terra30% ·
  4. GPT-5.4 Pro30% ·
  5. Claude Opus 529.1% ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1GPT-5.6 SolOpenAIAA public CritPt board, 1 Sep 2026. Headline leader on the evaluation page.2026-07-0932.3%
2026-09-01
Source-checked
2GPT-5.5 ProOpenAI2026-04-2330.6%
2026-09-01
Source-checked
3GPT-5.6 TerraOpenAI2026-07-0930%
2026-09-01
Source-checked
4GPT-5.4 ProOpenAI2026-03-0530%
2026-09-01
Source-checked
5Claude Opus 5Anthropic2026-07-2429.1%
2026-09-01
Source-checked
6Claude Fable 5Anthropic2026-06-0928.6%
2026-09-01
Source-checked
7GPT-5.5OpenAI2026-04-2327.1%
2026-09-01
Source-checked
8Gemini 3 Pro Deep ThinkGoogle2025-12-0125.7%
2026-09-01
Source-checked
9Kimi K3Moonshot2026-07-1623.4%
2026-09-01
Source-checked
11Claude Opus 4.8Anthropic2026-05-2820.9%
2026-09-01
Source-checked
15GLM-5.3Zhipu2026-08-1419.1%
2026-09-01
Source-checked
19Grok 4.6xAI2026-08-1217.1%
2026-09-01
Source-checked
Method

What it covers

Ledger rows are from Artificial Analysis' public CritPt evaluation (independently run). The 2025 paper reported ~4% best base-model accuracy; 2026 frontier rows sit near 30%. Dataset curation involved adversarial filtering against older models.

Score ceiling

Where it breaks down

Human expert time is research-scale. Published 2026 frontier accuracy clusters well below 40% on the composite set.

Tools that report it

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does CritPt measure?

Unpublished research-level physics: 71 composite challenges (70 test + 1 example) written by 50+ active physicists across 11 subfields, plus 190 simpler checkpoint tasks. 6% weight in AA Intelligence Index v4.1.

What does a high CritPt score not prove?

A strong CritPt result says nothing about:

  • Textbook or contest physics — problems are original research-scale tasks, not recycled qualifying-exam items.
  • Whether the derivation is a readable paper — scoring is machine-verifiable on guess-resistant formats (arrays, symbolic expressions, Python functions), not a proof grade.
  • Cross-domain scientific reasoning — CritPt is physics only. Use HLE or SciCode for broader science.

How is CritPt scored?

Average accuracy on the 70-test composite set. Automated grading for physics-specific output formats. VerdictPal rows use Artificial Analysis' independent CritPt run. Task format: 71 composite research challenges spanning condensed matter, quantum, AMO, astrophysics, HEP, mathematical physics, statistical physics, nuclear, nonlinear dynamics, fluids, and biophysics. Models may use coding tools in some AA runs; checkpoint tasks (190) are a separate, easier slice.

Is CritPt saturated?

CritPt is currently marked Active in the atlas. Ceiling context: Human expert time is research-scale. Published 2026 frontier accuracy clusters well below 40% on the composite set.

Can CritPt results be contaminated by training data?

Contamination risk for CritPt is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.