VerdictPal · editorial desk · 2026VerdictPal
Compare · Two models, head-to-head

Pick two models. See the diff.

Composite scores, benchmark ranks, capability grids, and pricing per million tokens. All sourced from the same data that powers the model atlas.

OpenAI o3 vs GPT-5.2

OpenAI o32stats won
Tied0stats
GPT-5.25stats won
Spec showdown

Capacity, price, freshness.

OpenAI o3vsGPT-5.2

Composite scores are within 5 points — 70 for OpenAI o3 vs 71 for GPT-5.2. The choice comes down to capability profile and price. GPT-5.2 has the higher Vals/AA intelligence score (42 vs 30). GPT-5.2's 410K context window dwarfs OpenAI o3's 200K — a meaningful difference for long-document or codebase work.

  1. Composite score
    OpenAI o370GPT-5.271Winner on this row
  2. Intelligence (Vals/AA)
    OpenAI o330GPT-5.242Winner on this row
  3. Context window
    OpenAI o3200KGPT-5.2410KWinner on this row
  4. Max output
    OpenAI o3100KGPT-5.2128KWinner on this row
  5. Input price / 1M (lower wins)
    OpenAI o3GPT-5.2$1.8Winner on this row
  6. Output speed (AA default API)
    OpenAI o3155 t/sWinner on this rowGPT-5.287 t/s
  7. Time to first token (AA default API)
    OpenAI o34.84sWinner on this rowGPT-5.297.23s
Composite score
70
vs
71Winner on this row
Intelligence (Vals/AA)
30
vs
42Winner on this row
Context window
200K
vs
410KWinner on this row
Max output
100K
vs
128KWinner on this row
Input price / 1M (lower wins)
vs
$1.8Winner on this row
Output speed (AA default API)
155 t/sWinner on this row
vs
87 t/s
Time to first token (AA default API)
4.84sWinner on this row
vs
97.23s
Capability grid

What each model can do.

Modalities

OpenAI o3GPT-5.2
Text
Image
Audio
Video
Code
Tool use
Reasoning

Access

OpenAI o3GPT-5.2
API
Consumer app
Open weights
Managed cloud
On-device
Family radar

Top score per benchmark family.

Knowledge
85
vs
87Winner on this row
Reasoning
88
vs
90Winner on this row
Math
99Winner on this row
vs
89
Coding
95
vs
95
Agentic
81
vs
85Winner on this row
Long context
69
vs
73Winner on this row
Benchmark ranks

Every shared benchmark, side by side.

Knowledge

MMLU-Pro#6/21 · #4/21
85.3%
vs
87.4%Winner on this row

Reasoning

Artificial Analysis Intelligence Index#48/67 · #29/67
30.4%
vs
42.2%Winner on this row
GPQA Diamond#43/69 · #22/69
82.7%
vs
90.3%Winner on this row
Humanity's Last Exam#45/65 · #27/65
20%
vs
35.4%Winner on this row
IFBench#30/57 · #19/57
71.43%
vs
75.44%Winner on this row

Math

AIME#8/19 · #6/19
88.33%
vs
89.2%Winner on this row
MATH#1/20 · #7/20
99.2%Winner on this row
vs
87.2%

Coding

LiveCodeBench#4/13 · #2/13
80.8%
vs
88.9%Winner on this row
SciCode#49/65 · #21/65
41%
vs
52.1%Winner on this row

Agentic

Terminal-Bench#55/69 · #46/69
37.12%
vs
46.97%Winner on this row
τ³-Bench Banking#11/64 · #9/64
80.7%
vs
84.8%Winner on this row

Long context

AA-LCR#27/64 · #12/64
69.33%
vs
72.67%Winner on this row

Performance

AA output speed#16/69 · #30/69
155 t/sWinner on this row
vs
87 t/s
AA time to first token#26/69 · #3/69
4.84sWinner on this row
vs
97.23s
Dossier facts

Dates, sources, fineprint.

OpenAI o3GPT-5.2
ProviderOpenAIOpenAI
FamilyOpenAI oGPT-5
Release date2025-04-162025-12-11
Knowledge cutoff2024-06-012025-08-01
Context window200K410K
Max output100K128K
Input price /1M$1.8
Output price /1M$14
Statussolidsolid
Composite score7071
Intelligence score3042
Intel sourceaaaa
Benchmark rows2118
Source count75
Freshnessfreshfresh
Pricing checked2026-06-092026-06-02
Ecosystem

Which tools wrap each model.

GPT-5.2

Not listed in any tool yet.

Verdict

Which model to actually pick.

Composite scores are within 5 points — 70 for OpenAI o3 vs 71 for GPT-5.2. The choice comes down to capability profile and price. GPT-5.2 has the higher Vals/AA intelligence score (42 vs 30). GPT-5.2's 410K context window dwarfs OpenAI o3's 200K — a meaningful difference for long-document or codebase work.