Models / OpenAI

GPT 5.5

GPT 5 · Released 2026-04-23

OpenAI's April 2026 flagship. Vals Index 67.62% (#2), AA Index 60.2 (#2), Terminal-Bench 2.0 82.0%, ARC-AGI-2 85% (top). LMArena Elo 1473 since April 2026. 512K context, $12.50 / $50 per 1M tokens.

Best generalist for coding agents, research workflows, and tool-heavy automation. The model the desk reaches for when we need coding agents that close the loop on a single issue.

#6 of 21 on Vals Index · current snapshot
68 Vals Index · 28 benchmark rows · 14 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite75
Context512K
Input / 1M$13
Output / 1M$50
Knowledge cutoff2026-03-15
Output speed90 t/s
TTFT (default API)31.38s
Statusflagship
KnowledgeReasoningMathCodingAgenticMultimodalLong contextTool useSafetyHuman preference
Knowledge
91
Reasoning
85
Math
92
Coding
96
Agentic
100
Multimodal
72
Long context
56
Tool use
87
Safety
91
Human preference
95

87average across 10 tested families

Editor's note
GPT 5.5 is the model the desk reaches for when we need coding agents that close the loop on a single issue. Terminal-Bench 2.0 (82.0%) and SWE-bench Verified (82.5%) are reproducible numbers, not vendor slideshows. The LMArena Elo lead is the most-cited number on the public market.
Benchmark placements

Where GPT 5.5 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
LMArena (Chatbot Arena)Human preference#1/ 1395LMArena leaderboard2026-05-25
ARC-AGIReasoning#2/ 1485ARC Prize leaderboard2026-05-15
BFCLTool use#2/ 2087Berkeley Function Calling Leaderboard V42026-04-15
DeepSWECoding#2/ 667DeepSWE leaderboard v1.12026-06-20
HumanEvalCoding#2/ 1596CodeSOTA HumanEval leaderboard2026-04-23
MMMUMultimodal#2/ 872MMMU leaderboard2026-04-23
GDPvalAgentic#3/ 6100Artificial Analysis — GDPval-AA v22026-06-17
HELM SafetySafety#3/ 991Stanford HELM leaderboard2026-04-23
LiveBenchReasoning#3/ 2170LiveBench leaderboard2026-04-15
LongBench v2Long context#3/ 1956LongBench v2 leaderboard2026-04-01
MCP AtlasTool use#3/ 1078MCP Atlas leaderboard2026-06-09
MMLUKnowledge#3/ 2291OpenAI model release notes2025-12-01
MATHMath#4/ 2090OpenAI model release notes2026-04-23
SWE-bench VerifiedCoding#4/ 2383SWE-bench Verified leaderboard2026-06-01
Vending-Bench 2Agentic#4/ 5Andon Labs Vending-Bench 22026-06-01
AIMEMath#5/ 1992OpenAI model release notes2026-04-23
Artificial Analysis Intelligence IndexReasoning#6/ 6755Artificial Analysis2026-07-21
Humanity's Last ExamReasoning#6/ 6544Artificial Analysis2026-07-21
Terminal-BenchAgentic#6/ 6982Terminal-Bench 2.0 leaderboard2026-05-20
Vals IndexAgentic#6/ 2068Vals AI — Vals Index2026-06-04
SciCodeCoding#7/ 6556Artificial Analysis2026-07-21
MMLU-ProKnowledge#10/ 2182MMLU-Pro HuggingFace leaderboard2026-04-23
AA time to first tokenPerformance#11/ 69Artificial Analysis2026-07-21
IFBenchReasoning#16/ 5776Artificial Analysis2026-07-21
τ³-Bench BankingAgentic#16/ 6455Artificial Analysis — τ³-Bench Banking2026-06-17
AA output speedPerformance#28/ 69Artificial Analysis2026-07-21
GPQA DiamondReasoning#47/ 6981GPQA Diamond paper leaderboard2026-05-15
AA-LCRLong context#54/ 6452Artificial Analysis — AA-LCR2026-06-17
Family context

The three highest-scoring carded models in each capability family. Where GPT 5.5 shows up, it's highlighted.

KnowledgeGPT 5.5 · 91
  1. 1
    91
  2. 2
    91
  3. 3
    91
ReasoningGPT 5.5 · 85
  1. 1
    95
  2. 2
    94
  3. 3
    94
MathGPT 5.5 · 92
  1. 1
    99
  2. 2
    97
  3. 3
    96
CodingGPT 5.5 · 96
  1. 1
    96
  2. 2
    96
  3. 3
    96
AgenticGPT 5.5 · 100
  1. 1
    100
  2. 2
    100
  3. 3
    100
MultimodalGPT 5.5 · 72
  1. 1
    72
  2. 2
    72
  3. 3
    71
Long contextGPT 5.5 · 56
  1. 1
    75
  2. 2
    75
  3. 3
    74
Tool useGPT 5.5 · 87
  1. 1
    89
  2. 2
    87
  3. 3
    87
Human preferenceGPT 5.5 · 95
  1. 1
    95
  2. 2
    92
  3. 3
    91
Newest receipts

Benchmark rows added to the public ledger for GPT 5.5 in the last 120 days. Older rows live in the full table below.

27 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Knowledge91
Reasoning85
Math92
Coding96
Agentic100
Multimodal72
Long context56
Tool use87
Safety91
Human preference95
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMiniMaxxAIMetaZhipuXiaomiMeituanCohere

Full benchmark ledger

Every catalog benchmark for GPT 5.5. Scores link to the original source; gaps mean no public row exists yet.

28 sourced rows

28 of 32 catalog benchmarks have a sourced row for GPT 5.5.

Knowledge

Reasoning

Math

Coding

Agentic

Multimodal

Long context

Tool use

Performance

Safety

Human preference

Changelog
  1. Released GPT 5.5 with native audio, image, and 512K context.
  2. GPT 5.4 retired for general availability.
Sources