Models / OpenAI

OpenAI o3

OpenAI o · Released 2025-04-16

OpenAI's April 2025 reasoning model. Extended thinking for math, code, and science. LiveBench 68.5%, LongBench v2 51.2%, BFCL 78.1%.

Reasoning-tier OpenAI pick when latency and cost are secondary to hardest single-shot problem solving.

#34 of 51 on AA Index · current snapshot
30 AA Index · 21 benchmark rows · 7 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite70
Context200K
Input / 1M
Output / 1M
Knowledge cutoff2024-06-01
Output speed155 t/s
TTFT (default API)4.84s
Statussolid
KnowledgeReasoningMathCodingAgenticLong contextTool use
Knowledge
85
Reasoning
88
Math
99
Coding
95
Agentic
81
Long context
69
Tool use
78

85average across 7 tested families

Benchmark placements

Where OpenAI o3 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
MATHMath#1/ 2099Artificial Analysis2026-07-21
LiveCodeBenchCoding#4/ 1381Artificial Analysis2026-07-21
MMLU-ProKnowledge#6/ 2185Artificial Analysis2026-07-21
AIMEMath#8/ 1988Artificial Analysis2026-07-21
τ³-Bench BankingAgentic#11/ 6481Artificial Analysis2026-07-21
AA output speedPerformance#16/ 69Artificial Analysis2026-07-21
AA time to first tokenPerformance#26/ 69Artificial Analysis2026-07-21
AA-LCRLong context#27/ 6469Artificial Analysis2026-07-21
IFBenchReasoning#30/ 5771Artificial Analysis2026-07-21
GPQA DiamondReasoning#43/ 6983Artificial Analysis2026-07-21
Humanity's Last ExamReasoning#45/ 6520Artificial Analysis2026-07-21
Artificial Analysis Intelligence IndexReasoning#48/ 6730Artificial Analysis2026-07-21
SciCodeCoding#49/ 6541Artificial Analysis2026-07-21
Terminal-BenchAgentic#55/ 6937Artificial Analysis2026-07-21
Family context

The three highest-scoring carded models in each capability family. Where OpenAI o3 shows up, it's highlighted.

KnowledgeOpenAI o3 · 85
  1. 1
    91
  2. 2
    91
  3. 3
    91
ReasoningOpenAI o3 · 88
  1. 1
    95
  2. 2
    94
  3. 3
    94
MathOpenAI o3 · 99
  1. 1
    99
  2. 2
    97
  3. 3
    96
CodingOpenAI o3 · 95
  1. 1
    96
  2. 2
    96
  3. 3
    96
AgenticOpenAI o3 · 81
  1. 1
    100
  2. 2
    100
  3. 3
    100
Long contextOpenAI o3 · 69
  1. 1
    75
  2. 2
    75
  3. 3
    74
Tool useOpenAI o3 · 78
  1. 1
    89
  2. 2
    87
  3. 3
    87
Newest receipts

Benchmark rows added to the public ledger for OpenAI o3 in the last 120 days. Older rows live in the full table below.

17 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Knowledge85
Reasoning88
Math99
Coding95
Agentic81
Long context69
Tool use78
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMiniMaxxAIMetaZhipuXiaomiMeituanCohere

Full benchmark ledger

Every catalog benchmark for OpenAI o3. Scores link to the original source; gaps mean no public row exists yet.

21 sourced rows

19 of 32 catalog benchmarks have a sourced row for OpenAI o3.

Knowledge

Reasoning

Math

Coding

Agentic

Long context

Tool use

Performance

Sources