VerdictPal · editorial desk · 2026VerdictPal
Compare · Two models, head-to-head

Pick two models. See the diff.

Composite scores, benchmark ranks, capability grids, and pricing per million tokens. All sourced from the same data that powers the model atlas.

Claude Sonnet 4.6 vs GPT 5.4

Claude Sonnet 4.66stats won
Tied0stats
GPT 5.41stats won
Spec showdown

Capacity, price, freshness.

Claude Sonnet 4.6vsGPT 5.4

Composite scores are within 5 points — 66 for Claude Sonnet 4.6 vs 61 for GPT 5.4. The choice comes down to capability profile and price. Claude Sonnet 4.6 has the higher Vals/AA intelligence score (60 vs 51). Claude Sonnet 4.6 is more affordable at $3.0/1M input — about 2× cheaper than $7.0/1M. Claude Sonnet 4.6's 1M context window dwarfs GPT 5.4's 256K — a meaningful difference for long-document or codebase work.

  1. Composite score
    Claude Sonnet 4.666Winner on this rowGPT 5.461
  2. Intelligence (Vals/AA)
    Claude Sonnet 4.660Winner on this rowGPT 5.451
  3. Context window
    Claude Sonnet 4.61MWinner on this rowGPT 5.4256K
  4. Max output
    Claude Sonnet 4.6128KWinner on this rowGPT 5.466K
  5. Input price / 1M (lower wins)
    Claude Sonnet 4.6$3.0Winner on this rowGPT 5.4$7.0
  6. Output speed (AA default API)
    Claude Sonnet 4.660 t/sGPT 5.4152 t/sWinner on this row
  7. Time to first token (AA default API)
    Claude Sonnet 4.61.18sWinner on this rowGPT 5.4129.43s
Composite score
66Winner on this row
vs
61
Intelligence (Vals/AA)
60Winner on this row
vs
51
Context window
1MWinner on this row
vs
256K
Max output
128KWinner on this row
vs
66K
Input price / 1M (lower wins)
$3.0Winner on this row
vs
$7.0
Output speed (AA default API)
60 t/s
vs
152 t/sWinner on this row
Time to first token (AA default API)
1.18sWinner on this row
vs
129.43s
Capability grid

What each model can do.

Modalities

Claude Sonnet 4.6GPT 5.4
Text
Image
Audio
Video
Code
Tool use
Reasoning

Access

Claude Sonnet 4.6GPT 5.4
API
Consumer app
Open weights
Managed cloud
On-device
Family radar

Top score per benchmark family.

Knowledge
91Winner on this row
vs
90
Reasoning
77
vs
92Winner on this row
Coding
94Winner on this row
vs
80
Agentic
80
vs
82Winner on this row
Long context
58
vs
74Winner on this row
Tool use
89Winner on this row
vs
84
Human preference
79
vs
81Winner on this row
Benchmark ranks

Every shared benchmark, side by side.

Knowledge

MMLU#4/22 · #7/22
90.6%Winner on this row
vs
89.8%

Reasoning

ARC-AGI#8/14 · #14/14
58.3%Winner on this row
vs
0.2%
Artificial Analysis Intelligence Index#43/67 · #11/67
35.9%
vs
51.4%Winner on this row
GPQA Diamond#54/69 · #13/69
77.2%
vs
92%Winner on this row
Humanity's Last Exam#54/65 · #10/65
13.2%
vs
41.6%Winner on this row
IFBench#54/57 · #22/57
41.16%
vs
73.95%Winner on this row
LiveBench#5/21 · #6/21
67.2%Winner on this row
vs
66.4%

Coding

SciCode#37/65 · #5/65
46.9%
vs
56.6%Winner on this row
SWE-bench Verified#13/23 · #12/23
79.6%
vs
80%Winner on this row

Agentic

Terminal-Bench#24/69 · #7/69
68.4%
vs
81.8%Winner on this row
τ³-Bench Banking#12/64 · #26/64
79.53%Winner on this row
vs
30.31%

Long context

AA-LCR#51/64 · #3/64
57.67%
vs
74%Winner on this row
LongBench v2#8/19 · #6/19
52.6%
vs
54.2%Winner on this row

Tool use

BFCL#1/20 · #5/20
88.7%Winner on this row
vs
83.8%

Human preference

LMArena (Chatbot Arena)#10/13 · #8/13
1396
vs
1405Winner on this row

Performance

AA output speed#48/69 · #17/69
60 t/s
vs
152 t/sWinner on this row
AA time to first token#39/69 · #1/69
1.18sWinner on this row
vs
129.43s
Dossier facts

Dates, sources, fineprint.

Claude Sonnet 4.6GPT 5.4
ProviderAnthropicOpenAI
FamilyClaude SonnetGPT 5
Release date2026-02-172026-03-05
Knowledge cutoff2025-12-012026-01-31
Context window1M256K
Max output128K66K
Input price /1M$3.0$7.0
Output price /1M$15$28
Statusflagshipflagship
Composite score6661
Intelligence score6051
Intel sourcevalsaa
Benchmark rows2418
Source count1110
Freshnessfreshfresh
Pricing checked2026-06-012026-06-01
Ecosystem

Which tools wrap each model.

Verdict

Which model to actually pick.

Composite scores are within 5 points — 66 for Claude Sonnet 4.6 vs 61 for GPT 5.4. The choice comes down to capability profile and price. Claude Sonnet 4.6 has the higher Vals/AA intelligence score (60 vs 51). Claude Sonnet 4.6 is more affordable at $3.0/1M input — about 2× cheaper than $7.0/1M. Claude Sonnet 4.6's 1M context window dwarfs GPT 5.4's 256K — a meaningful difference for long-document or codebase work.