Modelle / OpenAI

OpenAI o3

OpenAI o · Release 2025-04-16

OpenAIs Reasoning-Modell April 2025. Extended Thinking für Mathe, Code und Wissenschaft. LiveBench 68,5 %, LongBench v2 51,2 %, BFCL 78,1 %.

OpenAI-Reasoning-Tier, wenn Latenz und Kosten hinter schwierigster Single-Shot-Problemlösung zurückstehen.

#34 von 51 im AA Index · aktueller Snapshot
30 AA Index · 21 Benchmark-Zeilen · 7 QuellenPosition ist benchmark-spezifisch — kein Cross-Family- oder Cross-Source-Ranking.
Composite70
Kontext200K
Input / 1M
Output / 1M
Wissensstichtag2024-06-01
Output-Speed155 t/s
TTFT (Default-API)4.84s
Statussolid
KnowledgeReasoningMathCodingAgenticLong contextTool use
Knowledge
85
Reasoning
88
Math
99
Coding
95
Agentic
81
Long context
69
Tool use
78

85average across 7 tested families

Benchmark-Platzierungen

Wo OpenAI o3 auf jeder öffentlichen Benchmark-Quelle mit Zeile steht. Rang zählt jedes Modell mit neuester Zeile im selben Test — kein universeller Qualitätsscore.

BenchmarkFamilyRangScoreQuelleDatum
MATHMath#1/ 2099Artificial Analysis2026-07-21
LiveCodeBenchCoding#4/ 1381Artificial Analysis2026-07-21
MMLU-ProKnowledge#6/ 2185Artificial Analysis2026-07-21
AIMEMath#8/ 1988Artificial Analysis2026-07-21
τ³-Bench BankingAgentic#11/ 6481Artificial Analysis2026-07-21
AA output speedPerformance#16/ 69Artificial Analysis2026-07-21
AA time to first tokenPerformance#26/ 69Artificial Analysis2026-07-21
AA-LCRLong context#27/ 6469Artificial Analysis2026-07-21
IFBenchReasoning#30/ 5771Artificial Analysis2026-07-21
GPQA DiamondReasoning#43/ 6983Artificial Analysis2026-07-21
Humanity's Last ExamReasoning#45/ 6520Artificial Analysis2026-07-21
Artificial Analysis Intelligence IndexReasoning#48/ 6730Artificial Analysis2026-07-21
SciCodeCoding#49/ 6541Artificial Analysis2026-07-21
Terminal-BenchAgentic#55/ 6937Artificial Analysis2026-07-21
Family-Kontext

Die drei höchstscorierenden kartierten Modelle pro Capability-Family. Wo OpenAI o3 auftaucht, ist es hervorgehoben.

KnowledgeOpenAI o3 · 85
  1. 1
    91
  2. 2
    91
  3. 3
    91
ReasoningOpenAI o3 · 88
  1. 1
    95
  2. 2
    94
  3. 3
    94
MathOpenAI o3 · 99
  1. 1
    99
  2. 2
    97
  3. 3
    96
CodingOpenAI o3 · 95
  1. 1
    96
  2. 2
    96
  3. 3
    96
AgenticOpenAI o3 · 81
  1. 1
    100
  2. 2
    100
  3. 3
    100
Long contextOpenAI o3 · 69
  1. 1
    75
  2. 2
    75
  3. 3
    74
Tool useOpenAI o3 · 78
  1. 1
    89
  2. 2
    87
  3. 3
    87
Neueste Belege

Benchmark-Zeilen im öffentlichen Ledger für OpenAI o3 in den letzten 120 Tagen. Ältere Zeilen stehen in der vollen Tabelle unten.

17 neu
Family-Abdeckung

Ein hoher Composite, der eine schwache Family versteckt, ist eine Falle. Diese Balken zeigen Families ohne öffentlichen Test — und wo das Modell führt.

Knowledge85
Reasoning88
Math99
Coding95
Agentic81
Long context69
Tool use78
Preis vs. Performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMiniMaxxAIMetaZhipuXiaomiMeituanCohere

Volles Benchmark-Ledger

Jeder Katalog-Benchmark für OpenAI o3. Scores verlinken zur Originalquelle; Lücken heißen: noch keine öffentliche Zeile.

21 Quell-Zeilen

19 of 32 catalog benchmarks have a sourced row for OpenAI o3.

Knowledge

Reasoning

Math

Coding

Agentic

Long context

Tool use

Performance

Quellen