Models / OpenAI

GPT 5.4

GPT 5 · Released 2026-03-05

OpenAI's March 2026 model. AA Intelligence Index 56.8 (#5 globally). SWE-bench Verified 80.0%, Terminal-Bench 2.0 81.8% (ForgeCode). ARC-AGI-2 83.3% (Pro variant). 256K context, $7 / $28 per 1M tokens — 44% cheaper than GPT 5.5 on input.

Default for production OpenAI workloads where GPT 5.5 is overkill but GPT 5.2 is too old. The Pro variant leads ARC-AGI-2 among non-OpenAI competitors at 83.3%.

#14 of 168 on AA Index · current snapshot
53 AA Index · 19 benchmark rows · 10 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite62
Context256K
Input / 1M$7.0
Output / 1M$28
Knowledge cutoff2026-01-31
Statusflagship

Family profile

Best published score in each covered benchmark family.

79/100 avg
Knowledge
90
Reasoning
92
Math
48
Coding
80
Agentic
82
Long context
78
Tool use
84
Human preference
81

8 tested benchmark families

Editor's note
GPT 5.4 is the OpenAI model we recommend when GPT 5.5 is too expensive but the task still needs frontier coding. The ForgeCode variant of Terminal-Bench 2.0 (81.8%) is the most reproducible number we have on the model.
Benchmark placements

Where GPT 5.4 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
FrontierMathMath#4/ 1548Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15
BFCLTool use#5/ 2084Berkeley Function Calling Leaderboard V42026-04-15
DeepSWECoding#6/ 852DeepSWE leaderboard v1.12026-06-20
LiveBenchReasoning#6/ 2166LiveBench leaderboard2026-04-15
LongBench v2Long context#6/ 1954LongBench v2 leaderboard2026-04-01
SciCodeCoding#6/ 18057Artificial Analysis2026-09-01
MMLUKnowledge#7/ 2290OpenAI GPT-5.4 release2026-03-05
LMArena (Chatbot Arena)Human preference#8/ 1381LMArena leaderboard2026-05-25
SWE-bench VerifiedCoding#12/ 2380SWE-bench official leaderboard2026-05-30
ARC-AGIReasoning#14/ 140ARC-AGI-3 paper (arXiv 2603.24621)2026-03-20
Humanity's Last ExamReasoning#15/ 18044Artificial Analysis2026-09-01
τ³-Bench BankingAgentic#18/ 10640Artificial Analysis2026-09-01
Terminal-BenchAgentic#18/ 18482Terminal-Bench 2.0 leaderboard2026-05-20
AA-LCRLong context#21/ 17978Artificial Analysis2026-09-01
GPQA DiamondReasoning#23/ 18592Artificial Analysis2026-09-01
Artificial Analysis Intelligence IndexReasoning#24/ 18753Artificial Analysis2026-09-01
IFBenchReasoning#32/ 12774Artificial Analysis2026-09-01
AA time to first tokenPerformance#123/ 186Artificial Analysis2026-09-01
AA output speedPerformance#124/ 187Artificial Analysis2026-09-01
Family context

The three highest-scoring models with pages in each capability family. Where GPT 5.4 shows up, it's highlighted.

Newest receipts

Benchmark rows added to the public ledger for GPT 5.4 in the last 120 days. Older rows live in the full table below.

14 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Knowledge90
Reasoning92
Math48
Coding80
Agentic82
Long context78
Tool use84
Human preference81
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover any dot for name, score, and input price. Keyboard: tab through the top twelve, or use the ranking below.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMetaMiniMaxxAIZhipuXiaomiOtherMeituanCohereNVIDIA

Full benchmark ledger

Every catalog benchmark for GPT 5.4. Scores link to the original source; gaps mean no public row exists yet.

19 sourced rows

19 of 37 catalog benchmarks have a sourced row for GPT 5.4.

Knowledge

Reasoning

Math

Coding

Agentic

Long context

Tool use

Performance

Human preference

Changelog
  1. Released GPT 5.4 — AA Index 56.8, Terminal-Bench 2.0 81.8% (ForgeCode).
  2. GPT 5.2 retired for general availability.
Sources