Models / OpenAI

OpenAI o1

OpenAI o · Released 2024-09-12

OpenAI's September 2024 reasoning model. AIME 2024 83.3% with cons@64 sampling. Predecessor to o3 and GPT-5 reasoning tiers.

Legacy reasoning tier — prefer o3 or GPT-5.5 for new math and code work.

#140 of 168 on AA Index · current snapshot
24 AA Index · 13 benchmark rows · 2 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite55
Context200K
Input / 1M
Output / 1M
Knowledge cutoff2023-10-01
Statuslightweight

Family profile

Best published score in each covered benchmark family.

67/100 avg
Knowledge
84
Reasoning
75
Math
97
Coding
68
Agentic
13
Long context
63

6 tested benchmark families

Benchmark placements

Where OpenAI o1 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
MATHMath#7/ 2597Artificial Analysis2026-09-01
LiveCodeBenchCoding#24/ 3168Artificial Analysis2026-09-01
AIMEMath#26/ 3783OpenAI: Learning to Reason2024-09-12
MMLU-ProKnowledge#26/ 3984Artificial Analysis2026-09-01
IFBenchReasoning#53/ 12770Artificial Analysis2026-09-01
AA-LCRLong context#129/ 17963Artificial Analysis2026-09-01
SciCodeCoding#154/ 18036Artificial Analysis2026-09-01
GPQA DiamondReasoning#156/ 18575Artificial Analysis2026-09-01
Artificial Analysis Intelligence IndexReasoning#159/ 18724Artificial Analysis2026-09-01
Humanity's Last ExamReasoning#163/ 1807Artificial Analysis2026-09-01
Terminal-BenchAgentic#173/ 18413Artificial Analysis2026-09-01
AA time to first tokenPerformance#180/ 186Artificial Analysis2026-09-01
AA output speedPerformance#181/ 187Artificial Analysis2026-09-01
Family context

The three highest-scoring models with pages in each capability family. Where OpenAI o1 shows up, it's highlighted.

Newest receipts

Benchmark rows added to the public ledger for OpenAI o1 in the last 120 days. Older rows live in the full table below.

12 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Knowledge84
Reasoning75
Math97
Coding68
Agentic13
Long context63
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover any dot for name, score, and input price. Keyboard: tab through the top twelve, or use the ranking below.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMetaMiniMaxxAIZhipuXiaomiOtherMeituanCohereNVIDIA

Full benchmark ledger

Every catalog benchmark for OpenAI o1. Scores link to the original source; gaps mean no public row exists yet.

13 sourced rows

13 of 37 catalog benchmarks have a sourced row for OpenAI o1.

Knowledge

Reasoning

Math

Coding

Agentic

Long context

Performance

Sources