OpenAI's April 2026 flagship. Vals Index 67.62% (#2), AA Index 60.2 (#2), Terminal-Bench 2.0 82.0%, ARC-AGI-2 85% (top). LMArena Elo 1473 since April 2026. 512K context, $12.50 / $50 per 1M tokens.
Best generalist for coding agents, research workflows, and tool-heavy automation. The model the desk reaches for when we need coding agents that close the loop on a single issue.
#6 of 24 on Vals Index · current snapshot
68 Vals Index · 31 benchmark rows · 14 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Best published score in each covered benchmark family.
87/100 avg
050100
Knowledge
91
Reasoning
85
Math
92
Coding
96
Agentic
100
Multimodal
72
Long context
56
Tool use
87
Safety
91
Human preference
95
10 tested benchmark families
Editor's note
GPT 5.5 is the model the desk reaches for when we need coding agents that close the loop on a single issue. Terminal-Bench 2.0 (82.0%) and SWE-bench Verified (82.5%) are reproducible numbers, not vendor slideshows. The LMArena Elo lead is the most-cited number on the public market.
Benchmark placements
Per-benchmark positions
Where GPT 5.5 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.