90.8%
MMLU
Models / OpenAI
GPT 5 · Released 2026-04-23
OpenAI's April 2026 flagship. Vals Index 67.62% (#2), AA Index 60.2 (#2), Terminal-Bench 2.0 82.0%, ARC-AGI-2 85% (top). LMArena Elo 1473 since April 2026. 512K context, $12.50 / $50 per 1M tokens.
Best generalist for coding agents, research workflows, and tool-heavy automation. The model the desk reaches for when we need coding agents that close the loop on a single issue.
87average across 10 tested families
Where GPT 5.5 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.
The three highest-scoring carded models in each capability family. Where GPT 5.5 shows up, it's highlighted.
Benchmark rows added to the public ledger for GPT 5.5 in the last 120 days. Older rows live in the full table below.
A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.
Hover or tab any dot for name, score, and input price.
Every catalog benchmark for GPT 5.5. Scores link to the original source; gaps mean no public row exists yet.
28 of 32 catalog benchmarks have a sourced row for GPT 5.5.
85%
ARC-AGI-2 semi-private
54.8%
AA Intelligence Index v4.1
81.2%
GPQA Diamond accuracy (CoT)
44.3%
HLE (AA run)
75.85%
IFBench (AA run)
69.8%
LiveBench
67%
DeepSWE pass@1 (xhigh effort)
96.2% pass@1
HumanEval
56.1%
SciCode (AA run)
82.5%
SWE-bench Verified
1531
GDPval-AA v2
82%
Terminal-Bench 2.0 (audited harness)
68%
Vals Index
7523.84
Final bank balance (USD)
54.8%
τ³-Bench Banking pass@1
71.8%
MMMU
52.1%
AA-LCR pass@1
56.1%
LongBench v2
90 t/s
Median output tokens/s (1k prompt, default provider)
31.38s
Median time to first token (1k prompt, default provider)
90.6%
HELM Safety
1473
LMArena Elo (Style Controlled)