Why this benchmark is useful
ARC Prize is the only benchmark series that repeatedly surfaces new inflection points — from reasoning systems (ARC-AGI-2) to agentic intelligence (ARC-AGI-3).
ARC-AGI
Fluid, sample-efficient reasoning across the ARC Prize benchmark series — from static grid puzzles (ARC-AGI-1/2) to interactive agentic environments (ARC-AGI-3). Built to resist memorization and reward genuine generalization.
ARC Prize is the only benchmark series that repeatedly surfaces new inflection points — from reasoning systems (ARC-AGI-2) to agentic intelligence (ARC-AGI-3).
Pick the edition first. ARC-AGI-2 scores are static puzzle accuracy; ARC-AGI-3 scores are environment solve-rates with sparse feedback. Never mix them in a composite.
March 2026 interactive benchmark: agents explore novel environments, acquire goals on the fly, and adapt without language instructions. Measures agentic fluid intelligence.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.6 (Max)Anthropic | 2026-03-09 | 0.5% | ARC-AGI-3 paper (arXiv 2603.24621)2026-03-20 | Source-checked |
| 2 | Gemini 3.1 Pro PreviewGoogle | 2026-02-19 | 0.4% | ARC-AGI-3 paper (arXiv 2603.24621)2026-03-20 | Source-checked |
| 3 | GPT 5.4 (High)OpenAI | 2026-04-15 | 0.2% | ARC-AGI-3 paper (arXiv 2603.24621)2026-03-20 | Source-checked |
| 4 | Grok 4.20 (Reasoning)xAI | 2026-03-09 | 0.1% | ARC-AGI-3 paper (arXiv 2603.24621)2026-03-20 | Source-checked |
This benchmark sits in the reasoning family. It uses Edition-dependent: grid-to-grid puzzles (v1/v2) or interactive turn-based environments (v3). Held-out private sets prevent overfitting. The score should travel with its task format, scoring method, source date, and benchmark version.
Humans solve 100% of environments. Official release rows clustered below 1%. OpenAI's 3 Sep 2026 GPT-6 Astra launch claims 99.9% — vendor-claimed, not an ARC Prize leaderboard row. Do not mark this edition defunct until the official board confirms it.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
Fluid, sample-efficient reasoning across the ARC Prize benchmark series — from static grid puzzles (ARC-AGI-1/2) to interactive agentic environments (ARC-AGI-3). Built to resist memorization and reward genuine generalization.
A strong ARC-AGI result says nothing about:
% of held-out tasks or environments solved. ARC-AGI-3 measures skill-acquisition efficiency over time. Task format: Edition-dependent: grid-to-grid puzzles (v1/v2) or interactive turn-based environments (v3). Held-out private sets prevent overfitting.
ARC-AGI is currently marked Active in the atlas. Ceiling context: Edition-dependent — switch versions above. ARC-AGI-3 remains the official unsolved edition until ARC Prize publishes a confirming leaderboard row.
Contamination risk for ARC-AGI is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.