Benchmarks / Agentic

GDPval

Real-work deliverables vs experts

GDPval-AA v2 — real knowledge-work deliverables across 44 occupations, graded by blind pairwise comparison against human experts (Elo anchored at 1000). 10% weight in AA Intelligence Index v4.2 (was 20% in v4.1).

What this does not measure
  • VerdictPal endorsement — GDPval is one source, with its own task-mix choice.
  • End-to-end autonomy — these are graded artifacts, not long-running autonomous workflows.
  • Domain coverage — the published task set is a sample, not a comprehensive census of economically valuable work.
Analysis

Why this benchmark is useful

MCQ suites miss whether a model can ship a real artifact. GDPval grades deliverables from 44 occupations against blind human experts — still a major Agents slice of AA Intelligence Index v4.2, now sharing that category with AA-Briefcase.

Scope

Coverage map

Task family
Agentic
Format
Deliverable-graded tasks from real occupations, scored against human expert baselines.
Scoring
Win-rate vs. human expert baseline (Elo-style, human = 1000). v2 raises turn limit to 250 and uses a rotating frontier-model judge panel.
Maintainer
OpenAI
Reading guide

How to read the scores

Read Elo-style win rates against a 1000-point human baseline, not raw percentages. v2 changes turn limits and the judge panel — do not compare to pre-v2 headlines.

Blind spots

What it does not cover

  • VerdictPal endorsement — GDPval is one source, with its own task-mix choice.
  • End-to-end autonomy — these are graded artifacts, not long-running autonomous workflows.
  • Domain coverage — the published task set is a sample, not a comprehensive census of economically valuable work.
Scores

Evidence ledger

6 rows
66 rows
1818best score
5source-checked
3sources
2026-06-17to 2025-12-11
3

api · api

Papers / model cards2

manual-snapshot · manual

Vendor claims1

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Fable 51818 ·
  2. Claude Opus 4.81638 ·
  3. GPT 5.5 (xhigh)1531 ·
  4. GPT 5.2 Pro74.1% · OpenAI GDPval paper (arXiv)
  5. GPT 5.5 Pro82.3% · OpenAI GDPval grading portal

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Fable 5Anthropic2026-06-091818
2026-06-17
Source-checked
1GPT 5.2 ProOpenAI2026-05-1274.1%
OpenAI GDPval paper (arXiv)2026-05-12
Source-checked
2Claude Opus 4.8Anthropic2026-05-281638
2026-06-17
Source-checked
2GPT 5.5 ProOpenAI2026-04-2382.3%
OpenAI GDPval grading portal2026-04-23
Source-checked
3GPT 5.5 (xhigh)OpenAI2026-04-231531
2026-06-17
Source-checked
6GPT 5.2OpenAI2025-12-1171.8%
OpenAI GDPval2025-12-11
Needs audit
Method

What it covers

This benchmark sits in the agentic family. It uses Deliverable-graded tasks from real occupations, scored against human expert baselines. The score should travel with its task format, scoring method, source date, and benchmark version.

Score ceiling

Where it breaks down

Graded artifacts are not autonomous multi-day workflows. Strong GDPval does not prove safe deployment in regulated or client-facing settings.

Tools that report it

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does GDPval measure?

GDPval-AA v2 — real knowledge-work deliverables across 44 occupations, graded by blind pairwise comparison against human experts (Elo anchored at 1000). 10% weight in AA Intelligence Index v4.2 (was 20% in v4.1).

What does a high GDPval score not prove?

A strong GDPval result says nothing about:

  • VerdictPal endorsement — GDPval is one source, with its own task-mix choice.
  • End-to-end autonomy — these are graded artifacts, not long-running autonomous workflows.
  • Domain coverage — the published task set is a sample, not a comprehensive census of economically valuable work.

How is GDPval scored?

Win-rate vs. human expert baseline (Elo-style, human = 1000). v2 raises turn limit to 250 and uses a rotating frontier-model judge panel. Task format: Deliverable-graded tasks from real occupations, scored against human expert baselines.

Is GDPval saturated?

GDPval is currently marked Active in the atlas.

Can GDPval results be contaminated by training data?

Contamination risk for GDPval is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.