Models / OpenAI

GPT 5.4

GPT 5 · Released 2026-03-05

OpenAI's March 2026 model. AA Intelligence Index 56.8 (#5 globally). SWE-bench Verified 80.0%, Terminal-Bench 2.0 81.8% (ForgeCode). ARC-AGI-2 83.3% (Pro variant). 256K context, $7 / $28 per 1M tokens — 44% cheaper than GPT 5.5 on input.

Default for production OpenAI workloads where GPT 5.5 is overkill but GPT 5.2 is too old. The Pro variant leads ARC-AGI-2 among non-OpenAI competitors at 83.3%.

#5 of 51 on AA Index · current snapshot
51 AA Index · 18 benchmark rows · 10 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite61
Context256K
Input / 1M$7.0
Output / 1M$28
Knowledge cutoff2026-01-31
Output speed152 t/s
TTFT (default API)129.43s
Statusflagship
KnowledgeReasoningCodingAgenticLong contextTool useHuman preference
Knowledge
90
Reasoning
92
Coding
80
Agentic
82
Long context
74
Tool use
84
Human preference
81

83average across 7 tested families

Editor's note
GPT 5.4 is the OpenAI model we recommend when GPT 5.5 is too expensive but the task still needs frontier coding. The ForgeCode variant of Terminal-Bench 2.0 (81.8%) is the most reproducible number we have on the model.
Benchmark placements

Where GPT 5.4 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
AA time to first tokenPerformance#1/ 69Artificial Analysis2026-07-21
AA-LCRLong context#3/ 6474Artificial Analysis2026-07-21
DeepSWECoding#4/ 652DeepSWE leaderboard v1.12026-06-20
BFCLTool use#5/ 2084Berkeley Function Calling Leaderboard V42026-04-15
SciCodeCoding#5/ 6557Artificial Analysis2026-07-21
LiveBenchReasoning#6/ 2166LiveBench leaderboard2026-04-15
LongBench v2Long context#6/ 1954LongBench v2 leaderboard2026-04-01
MMLUKnowledge#7/ 2290OpenAI GPT-5.4 release2026-03-05
Terminal-BenchAgentic#7/ 6982Terminal-Bench 2.0 leaderboard2026-05-20
LMArena (Chatbot Arena)Human preference#8/ 1381LMArena leaderboard2026-05-25
Humanity's Last ExamReasoning#10/ 6542Artificial Analysis2026-07-21
Artificial Analysis Intelligence IndexReasoning#11/ 6751Artificial Analysis2026-07-21
SWE-bench VerifiedCoding#12/ 2380SWE-bench official leaderboard2026-05-30
GPQA DiamondReasoning#13/ 6992Artificial Analysis2026-07-21
ARC-AGIReasoning#14/ 140ARC-AGI-3 paper (arXiv 2603.24621)2026-03-20
AA output speedPerformance#17/ 69Artificial Analysis2026-07-21
IFBenchReasoning#22/ 5774Artificial Analysis2026-07-21
τ³-Bench BankingAgentic#26/ 6430Artificial Analysis2026-07-21
Family context

The three highest-scoring carded models in each capability family. Where GPT 5.4 shows up, it's highlighted.

KnowledgeGPT 5.4 · 90
  1. 1
    91
  2. 2
    91
  3. 3
    91
ReasoningGPT 5.4 · 92
  1. 1
    95
  2. 2
    94
  3. 3
    94
CodingGPT 5.4 · 80
  1. 1
    96
  2. 2
    96
  3. 3
    96
AgenticGPT 5.4 · 82
  1. 1
    100
  2. 2
    100
  3. 3
    100
Long contextGPT 5.4 · 74
  1. 1
    75
  2. 2
    75
  3. 3
    74
Tool useGPT 5.4 · 84
  1. 1
    89
  2. 2
    87
  3. 3
    87
Human preferenceGPT 5.4 · 81
  1. 1
    95
  2. 2
    92
  3. 3
    91
Newest receipts

Benchmark rows added to the public ledger for GPT 5.4 in the last 120 days. Older rows live in the full table below.

16 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Knowledge90
Reasoning92
Coding80
Agentic82
Long context74
Tool use84
Human preference81
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMiniMaxxAIMetaZhipuXiaomiMeituanCohere

Full benchmark ledger

Every catalog benchmark for GPT 5.4. Scores link to the original source; gaps mean no public row exists yet.

18 sourced rows

18 of 32 catalog benchmarks have a sourced row for GPT 5.4.

Knowledge

Reasoning

Coding

Agentic

Long context

Tool use

Performance

Human preference

Changelog
  1. Released GPT 5.4 — AA Index 56.8, Terminal-Bench 2.0 81.8% (ForgeCode).
  2. GPT 5.2 retired for general availability.
Sources