Benchmarks / Reasoning

Abstraction and Reasoning Corpus

ARC-AGI

Fluid, sample-efficient reasoning across the ARC Prize benchmark series — from static grid puzzles (ARC-AGI-1/2) to interactive agentic environments (ARC-AGI-3). Built to resist memorization and reward genuine generalization.

What this does not measure
  • Knowledge or language — tasks are deliberately knowledge-free, so scores say nothing about factual accuracy.
  • Cost-blind scores mislead — some high ARC-AGI-1 results used enormous per-task compute; always read the $/task footnote.
  • Cross-version comparison — ARC-AGI-1, ARC-AGI-2, and ARC-AGI-3 are different tasks with different ceilings. Use the version switcher.
Analysis

Why this benchmark is useful

ARC Prize is the only benchmark series that repeatedly surfaces new inflection points — from reasoning systems (ARC-AGI-2) to agentic intelligence (ARC-AGI-3).

Scope

Coverage map

Task family
Reasoning
Format
Edition-dependent: grid-to-grid puzzles (v1/v2) or interactive turn-based environments (v3). Held-out private sets prevent overfitting.
Scoring
% of held-out tasks or environments solved. ARC-AGI-3 measures skill-acquisition efficiency over time.
Maintainer
François Chollet / ARC Prize Foundation
Reading guide

How to read the scores

Pick the edition first. ARC-AGI-2 scores are static puzzle accuracy; ARC-AGI-3 scores are environment solve-rates with sparse feedback. Never mix them in a composite.

Blind spots

What it does not cover

  • Knowledge or language — tasks are deliberately knowledge-free, so scores say nothing about factual accuracy.
  • Cost-blind scores mislead — some high ARC-AGI-1 results used enormous per-task compute; always read the $/task footnote.
  • Cross-version comparison — ARC-AGI-1, ARC-AGI-2, and ARC-AGI-3 are different tasks with different ceilings. Use the version switcher.
Scores

Evidence ledger

March 2026 interactive benchmark: agents explore novel environments, acquire goals on the fly, and adapt without language instructions. Measures agentic fluid intelligence.

4 rows
44 rows
0.5%best score
100%human ceiling
99.5 ptsheadroom
4source-checked
1sources
2026-03-20source date
Papers / model cards4

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Opus 4.6 (Max)0.5% · ARC-AGI-3 paper (arXiv 2603.24621)
  2. Gemini 3.1 Pro Preview0.4% · ARC-AGI-3 paper (arXiv 2603.24621)
  3. GPT 5.4 (High)0.2% · ARC-AGI-3 paper (arXiv 2603.24621)
  4. Grok 4.20 (Reasoning)0.1% · ARC-AGI-3 paper (arXiv 2603.24621)

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Opus 4.6 (Max)Anthropic2026-03-090.5%
ARC-AGI-3 paper (arXiv 2603.24621)2026-03-20
Source-checked
2Gemini 3.1 Pro PreviewGoogle2026-02-190.4%
ARC-AGI-3 paper (arXiv 2603.24621)2026-03-20
Source-checked
3GPT 5.4 (High)OpenAI2026-04-150.2%
ARC-AGI-3 paper (arXiv 2603.24621)2026-03-20
Source-checked
4Grok 4.20 (Reasoning)xAI2026-03-090.1%
ARC-AGI-3 paper (arXiv 2603.24621)2026-03-20
Source-checked
Method

What it covers

This benchmark sits in the reasoning family. It uses Edition-dependent: grid-to-grid puzzles (v1/v2) or interactive turn-based environments (v3). Held-out private sets prevent overfitting. The score should travel with its task format, scoring method, source date, and benchmark version.

Score ceiling

Where it breaks down

Humans solve 100% of environments. Official release rows clustered below 1%. OpenAI's 3 Sep 2026 GPT-6 Astra launch claims 99.9% — vendor-claimed, not an ARC Prize leaderboard row. Do not mark this edition defunct until the official board confirms it.

Tools that report it

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does ARC-AGI measure?

Fluid, sample-efficient reasoning across the ARC Prize benchmark series — from static grid puzzles (ARC-AGI-1/2) to interactive agentic environments (ARC-AGI-3). Built to resist memorization and reward genuine generalization.

What does a high ARC-AGI score not prove?

A strong ARC-AGI result says nothing about:

  • Knowledge or language — tasks are deliberately knowledge-free, so scores say nothing about factual accuracy.
  • Cost-blind scores mislead — some high ARC-AGI-1 results used enormous per-task compute; always read the $/task footnote.
  • Cross-version comparison — ARC-AGI-1, ARC-AGI-2, and ARC-AGI-3 are different tasks with different ceilings. Use the version switcher.

How is ARC-AGI scored?

% of held-out tasks or environments solved. ARC-AGI-3 measures skill-acquisition efficiency over time. Task format: Edition-dependent: grid-to-grid puzzles (v1/v2) or interactive turn-based environments (v3). Held-out private sets prevent overfitting.

Is ARC-AGI saturated?

ARC-AGI is currently marked Active in the atlas. Ceiling context: Edition-dependent — switch versions above. ARC-AGI-3 remains the official unsolved edition until ARC Prize publishes a confirming leaderboard row.

Can ARC-AGI results be contaminated by training data?

Contamination risk for ARC-AGI is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.