Models / OpenAI

OpenAI o3

OpenAI o · Released 2025-04-16

OpenAI's April 2025 reasoning model. Extended thinking for math, code, and science. LiveBench 68.5%, LongBench v2 51.2%, BFCL 78.1%.

Reasoning-tier OpenAI pick when latency and cost are secondary to hardest single-shot problem solving.

#125 of 168 on AA Index · current snapshot
31 AA Index · 20 benchmark rows · 7 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite70
Context200K
Input / 1M
Output / 1M
Knowledge cutoff2024-06-01
Output speed134 t/s
TTFT (default API)5.65s
Statuslightweight

Family profile

Best published score in each covered benchmark family.

79/100 avg
Knowledge
85
Reasoning
88
Math
99
Coding
95
Agentic
37
Long context
73
Tool use
78

7 tested benchmark families

Benchmark placements

Where OpenAI o3 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
MATHMath#2/ 2599Artificial Analysis2026-09-01
LiveCodeBenchCoding#17/ 3181Artificial Analysis2026-09-01
AIMEMath#20/ 3788Artificial Analysis2026-09-01
AA time to first tokenPerformance#21/ 186Artificial Analysis2026-09-01
AA output speedPerformance#22/ 187Artificial Analysis2026-09-01
MMLU-ProKnowledge#22/ 3985Artificial Analysis2026-09-01
IFBenchReasoning#46/ 12771Artificial Analysis2026-09-01
AA-LCRLong context#62/ 17973Artificial Analysis2026-09-01
SciCodeCoding#116/ 18041Artificial Analysis2026-09-01
Humanity's Last ExamReasoning#120/ 18020Artificial Analysis2026-09-01
GPQA DiamondReasoning#124/ 18583Artificial Analysis2026-09-01
Terminal-BenchAgentic#126/ 18437Artificial Analysis2026-09-01
Artificial Analysis Intelligence IndexReasoning#143/ 18731Artificial Analysis2026-09-01
Family context

The three highest-scoring models with pages in each capability family. Where OpenAI o3 shows up, it's highlighted.

Newest receipts

Benchmark rows added to the public ledger for OpenAI o3 in the last 120 days. Older rows live in the full table below.

13 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Knowledge85
Reasoning88
Math99
Coding95
Agentic37
Long context73
Tool use78
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover any dot for name, score, and input price. Keyboard: tab through the top twelve, or use the ranking below.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMetaMiniMaxxAIZhipuXiaomiOtherMeituanCohereNVIDIA

Full benchmark ledger

Every catalog benchmark for OpenAI o3. Scores link to the original source; gaps mean no public row exists yet.

20 sourced rows

18 of 37 catalog benchmarks have a sourced row for OpenAI o3.

Knowledge

Reasoning

Math

Coding

Agentic

Long context

Tool use

Performance

Sources