Why this benchmark is useful
When GPQA and HLE no longer tell you whether a model can do research-scale physics: original, unpublished problems with guess-resistant answers, still far from ceiling (frontier around 30%).
Critical Physics Tasks
Unpublished research-level physics: 71 composite challenges (70 test + 1 example) written by 50+ active physicists across 11 subfields, plus 190 simpler checkpoint tasks. 6% weight in AA Intelligence Index v4.1.
When GPQA and HLE no longer tell you whether a model can do research-scale physics: original, unpublished problems with guess-resistant answers, still far from ceiling (frontier around 30%).
Read the composite challenge accuracy, not checkpoint-only headlines. GPT-5.6 Sol (max) leads the 1 Sep 2026 AA board at 32.3%. Tool-augmented runs are a separate, higher tier — do not mix them with base rows.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | GPT-5.6 SolOpenAIAA public CritPt board, 1 Sep 2026. Headline leader on the evaluation page. | 2026-07-09 | 32.3% | 2026-09-01 | Source-checked |
| 2 | GPT-5.5 ProOpenAI | 2026-04-23 | 30.6% | 2026-09-01 | Source-checked |
| 3 | GPT-5.6 TerraOpenAI | 2026-07-09 | 30% | 2026-09-01 | Source-checked |
| 4 | GPT-5.4 ProOpenAI | 2026-03-05 | 30% | 2026-09-01 | Source-checked |
| 5 | Claude Opus 5Anthropic | 2026-07-24 | 29.1% | 2026-09-01 | Source-checked |
| 6 | Claude Fable 5Anthropic | 2026-06-09 | 28.6% | 2026-09-01 | Source-checked |
| 7 | GPT-5.5OpenAI | 2026-04-23 | 27.1% | 2026-09-01 | Source-checked |
| 8 | Gemini 3 Pro Deep ThinkGoogle | 2025-12-01 | 25.7% | 2026-09-01 | Source-checked |
| 9 | Kimi K3Moonshot | 2026-07-16 | 23.4% | 2026-09-01 | Source-checked |
| 11 | Claude Opus 4.8Anthropic | 2026-05-28 | 20.9% | 2026-09-01 | Source-checked |
| 15 | GLM-5.3Zhipu | 2026-08-14 | 19.1% | 2026-09-01 | Source-checked |
| 19 | Grok 4.6xAI | 2026-08-12 | 17.1% | 2026-09-01 | Source-checked |
Ledger rows are from Artificial Analysis' public CritPt evaluation (independently run). The 2025 paper reported ~4% best base-model accuracy; 2026 frontier rows sit near 30%. Dataset curation involved adversarial filtering against older models.
Human expert time is research-scale. Published 2026 frontier accuracy clusters well below 40% on the composite set.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
Unpublished research-level physics: 71 composite challenges (70 test + 1 example) written by 50+ active physicists across 11 subfields, plus 190 simpler checkpoint tasks. 6% weight in AA Intelligence Index v4.1.
A strong CritPt result says nothing about:
Average accuracy on the 70-test composite set. Automated grading for physics-specific output formats. VerdictPal rows use Artificial Analysis' independent CritPt run. Task format: 71 composite research challenges spanning condensed matter, quantum, AMO, astrophysics, HEP, mathematical physics, statistical physics, nuclear, nonlinear dynamics, fluids, and biophysics. Models may use coding tools in some AA runs; checkpoint tasks (190) are a separate, easier slice.
CritPt is currently marked Active in the atlas. Ceiling context: Human expert time is research-scale. Published 2026 frontier accuracy clusters well below 40% on the composite set.
Contamination risk for CritPt is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.