Benchmarks / Agentic

GDPval

Real-work deliverables vs experts

GDPval-AA v2 — real knowledge-work deliverables across 44 occupations, graded by blind pairwise comparison against human experts (Elo anchored at 1000). 20% weight in AA Intelligence Index v4.1.

AgenticSolidActiveLow contamination riskSince 2025
What this does not measure
  • VerdictPal endorsement — GDPval is one source, with its own task-mix choice.
  • End-to-end autonomy — these are graded artifacts, not long-running autonomous workflows.
  • Domain coverage — the published task set is a sample, not a comprehensive census of economically valuable work.
Analysis

Why this benchmark is useful

MCQ suites miss whether a model can ship a real artifact. GDPval grades deliverables from 44 occupations against blind human experts — the heaviest slice of AA Intelligence Index v4.1.

Scope

Coverage map

Task family
Agentic
Format
Deliverable-graded tasks from real occupations, scored against human expert baselines.
Scoring
Win-rate vs. human expert baseline (Elo-style, human = 1000). v2 raises turn limit to 250 and uses a rotating frontier-model judge panel.
Maintainer
OpenAI
Reading guide

How to read the scores

Read Elo-style win rates against a 1000-point human baseline, not raw percentages. v2 changes turn limits and the judge panel — do not compare to pre-v2 headlines.

Blind spots

What it does not cover

  • VerdictPal endorsement — GDPval is one source, with its own task-mix choice.
  • End-to-end autonomy — these are graded artifacts, not long-running autonomous workflows.
  • Domain coverage — the published task set is a sample, not a comprehensive census of economically valuable work.
Scores

Evidence ledger

6 rows
66 rows
1818best score
5source-checked
3sources
2026-06-17to 2025-12-11
3

api · api

Papers / model cards2

manual-snapshot · manual

Vendor claims1

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Fable 51818 ·
  2. Claude Opus 4.81638 ·
  3. GPT 5.5 (xhigh)1531 ·
  4. GPT 5.2 Pro74.1% · OpenAI GDPval paper (arXiv)
  5. GPT 5.5 Pro82.3% · OpenAI GDPval grading portal

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Fable 5Anthropic2026-06-091818
2026-06-17
Source-checked
1GPT 5.2 ProOpenAI2026-05-1274.1%
OpenAI GDPval paper (arXiv)2026-05-12
Source-checked
2Claude Opus 4.8Anthropic2026-05-281638
2026-06-17
Source-checked
2GPT 5.5 ProOpenAI2026-04-2382.3%
OpenAI GDPval grading portal2026-04-23
Source-checked
3GPT 5.5 (xhigh)OpenAI2026-04-231531
2026-06-17
Source-checked
6GPT 5.2OpenAI2025-12-1171.8%
OpenAI GDPval2025-12-11
Needs audit
Method

What it covers

This benchmark sits in the agentic family. It uses Deliverable-graded tasks from real occupations, scored against human expert baselines. The score should travel with its task format, scoring method, source date, and benchmark version.

Score ceiling

Where it breaks down

Graded artifacts are not autonomous multi-day workflows. Strong GDPval does not prove safe deployment in regulated or client-facing settings.

Tools that report it

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.