Benchmarks / Reasoning

Abstraction and Reasoning Corpus

ARC-AGI

Fluid, sample-efficient reasoning across the ARC Prize benchmark series — from static grid puzzles (ARC-AGI-1/2) to interactive agentic environments (ARC-AGI-3). Built to resist memorization and reward genuine generalization.

ReasoningFlagshipActiveLow contamination riskSince 2019
What this does not measure
  • Knowledge or language — tasks are deliberately knowledge-free, so scores say nothing about factual accuracy.
  • Cost-blind scores mislead — some high ARC-AGI-1 results used enormous per-task compute; always read the $/task footnote.
  • Cross-version comparison — ARC-AGI-1, ARC-AGI-2, and ARC-AGI-3 are different tasks with different ceilings. Use the version switcher.
Analysis

Why this benchmark is useful

ARC Prize is the only benchmark series that repeatedly surfaces new inflection points — from reasoning systems (ARC-AGI-2) to agentic intelligence (ARC-AGI-3).

Scope

Coverage map

Task family
Reasoning
Format
Edition-dependent: grid-to-grid puzzles (v1/v2) or interactive turn-based environments (v3). Held-out private sets prevent overfitting.
Scoring
% of held-out tasks or environments solved. ARC-AGI-3 measures skill-acquisition efficiency over time.
Maintainer
François Chollet / ARC Prize Foundation
Reading guide

How to read the scores

Pick the edition first. ARC-AGI-2 scores are static puzzle accuracy; ARC-AGI-3 scores are environment solve-rates with sparse feedback. Never mix them in a composite.

Blind spots

What it does not cover

  • Knowledge or language — tasks are deliberately knowledge-free, so scores say nothing about factual accuracy.
  • Cost-blind scores mislead — some high ARC-AGI-1 results used enormous per-task compute; always read the $/task footnote.
  • Cross-version comparison — ARC-AGI-1, ARC-AGI-2, and ARC-AGI-3 are different tasks with different ceilings. Use the version switcher.
Scores

Evidence ledger

March 2026 interactive benchmark: agents explore novel environments, acquire goals on the fly, and adapt without language instructions. Measures agentic fluid intelligence.

4 rows
44 rows
0.5%best score
100%human ceiling
99.5 ptsheadroom
4source-checked
1sources
2026-03-20source date
Papers / model cards4

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Opus 4.6 (Max)0.5% · ARC-AGI-3 paper (arXiv 2603.24621)
  2. Gemini 3.1 Pro Preview0.4% · ARC-AGI-3 paper (arXiv 2603.24621)
  3. GPT 5.4 (High)0.2% · ARC-AGI-3 paper (arXiv 2603.24621)
  4. Grok 4.20 (Reasoning)0.1% · ARC-AGI-3 paper (arXiv 2603.24621)

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Opus 4.6 (Max)Anthropic2026-03-090.5%
ARC-AGI-3 paper (arXiv 2603.24621)2026-03-20
Source-checked
2Gemini 3.1 Pro PreviewGoogle2026-02-190.4%
ARC-AGI-3 paper (arXiv 2603.24621)2026-03-20
Source-checked
3GPT 5.4 (High)OpenAI2026-04-150.2%
ARC-AGI-3 paper (arXiv 2603.24621)2026-03-20
Source-checked
4Grok 4.20 (Reasoning)xAI2026-03-090.1%
ARC-AGI-3 paper (arXiv 2603.24621)2026-03-20
Source-checked
Method

What it covers

This benchmark sits in the reasoning family. It uses Edition-dependent: grid-to-grid puzzles (v1/v2) or interactive turn-based environments (v3). Held-out private sets prevent overfitting. The score should travel with its task format, scoring method, source date, and benchmark version.

Score ceiling

Where it breaks down

Humans solve 100% of environments. Frontier models score below 1% at release (Opus 4.6 ~0.5%, Gemini 3.1 Pro ~0.4%).

Tools that report it

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.