Models / Anthropic

Claude 4.1 Opus (Reasoning)

Claude 4.1 · Released 2025-08-05

AA Intelligence Index 33.7. Listed on Artificial Analysis with API pricing $15/$75 per 1M tokens. Benchmark rows sync from the AA snapshot — see the ledger on this page.

Frontier model tracked on Artificial Analysis. Use the benchmark ledger below for sourced scores; we do not infer numbers beyond published rows.

#78 of 131 on AA Index · current snapshot
34 AA Index · 13 benchmark rows · 1 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite55
ContextUnlisted
Input / 1M$15
Output / 1M$75
Knowledge cutoffunpublished
Output speed0 t/s
TTFT (default API)0.00s
Statussolid
KnowledgeReasoningMathCodingAgenticLong context
Knowledge
88
Reasoning
81
Math
80
Coding
65
Agentic
71
Long context
66

75average across 6 tested families

Benchmark placements

Where Claude 4.1 Opus (Reasoning) places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
MMLU-ProKnowledge#7/ 3488Artificial Analysis2026-07-28
LiveCodeBenchCoding#21/ 2665Artificial Analysis2026-07-28
AIMEMath#25/ 3280Artificial Analysis2026-07-28
τ³-Bench BankingAgentic#45/ 14371Artificial Analysis2026-07-28
AA-LCRLong context#55/ 14366Artificial Analysis2026-07-28
AA output speedPerformance#74/ 149Artificial Analysis2026-07-28
AA time to first tokenPerformance#74/ 149Artificial Analysis2026-07-28
IFBenchReasoning#87/ 12155Artificial Analysis2026-07-28
Artificial Analysis Intelligence IndexReasoning#91/ 14734Artificial Analysis2026-07-28
SciCodeCoding#95/ 14441Artificial Analysis2026-07-28
GPQA DiamondReasoning#107/ 14981Artificial Analysis2026-07-28
Terminal-BenchAgentic#112/ 14834Artificial Analysis2026-07-28
Humanity's Last ExamReasoning#117/ 14412Artificial Analysis2026-07-28
Family context

The three highest-scoring carded models in each capability family. Where Claude 4.1 Opus (Reasoning) shows up, it's highlighted.

KnowledgeClaude 4.1 Opus (Reasoning) · 88
  1. 1
    91
  2. 2
    91
  3. 3
    91
ReasoningClaude 4.1 Opus (Reasoning) · 81
  1. 1
    95
  2. 2
    94
  3. 3
    94
MathClaude 4.1 Opus (Reasoning) · 80
  1. 1
    99
  2. 2
    99
  3. 3
    99
CodingClaude 4.1 Opus (Reasoning) · 65
  1. 1
    96
  2. 2
    96
  3. 3
    96
AgenticClaude 4.1 Opus (Reasoning) · 71
  1. 1
    100
  2. 2
    100
  3. 3
    100
Long contextClaude 4.1 Opus (Reasoning) · 66
  1. 1
    76
  2. 2
    76
  3. 3
    75
Newest receipts

Benchmark rows added to the public ledger for Claude 4.1 Opus (Reasoning) in the last 120 days. Older rows live in the full table below.

13 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Knowledge88
Reasoning81
Math80
Coding65
Agentic71
Long context66
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotxAIMiniMaxMetaZhipuXiaomiOtherMeituanCohereNVIDIA

Full benchmark ledger

Every catalog benchmark for Claude 4.1 Opus (Reasoning). Scores link to the original source; gaps mean no public row exists yet.

13 sourced rows

13 of 32 catalog benchmarks have a sourced row for Claude 4.1 Opus (Reasoning).

Knowledge

Reasoning

Math

Coding

Agentic

Long context

Performance

Sources