Models / Anthropic

Claude Opus 4.8

Claude Opus · Released 2026-05-28

Anthropic's May 2026 flagship. Tops the Vals Index (70.17%) and the Artificial Analysis Intelligence Index (61.4). SWE-bench Verified 88.6%, Terminal-Bench 2.0 74.6%, 1M-token context, $15 / $75 per 1M tokens.

Front-door model for hard reasoning, long documents, agentic coding, and tool-heavy research. The model we reach for first when a task is too hard for Sonnet but small enough to fit in one prompt.

#5 of 21 on Vals Index · current snapshot
70 Vals Index · 26 benchmark rows · 12 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite73
Context1M
Input / 1M$15
Output / 1M$75
Knowledge cutoff2026-02-01
Output speed63 t/s
TTFT (default API)30.13s
Statusflagship
KnowledgeReasoningMathCodingAgenticMultimodalLong contextTool useSafetyHuman preference
Knowledge
91
Reasoning
80
Math
94
Coding
96
Agentic
100
Multimodal
71
Long context
58
Tool use
87
Safety
92
Human preference
92

86average across 10 tested families

Editor's note
Opus 4.8 is the model we reach for first when a task is hard enough to make a Sonnet hallucinate but small enough to fit in one prompt. The 1M context window is real, not marketing — long-document citation fidelity and BFCL scores back it up. The single-number rank (Vals, AA) is the highest we've measured.
Benchmark placements

Where Claude Opus 4.8 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
LongBench v2Long context#1/ 1958LongBench v2 leaderboard2026-04-01
MMLUKnowledge#1/ 2291Anthropic Claude 4 family2026-05-28
AIMEMath#2/ 1994Anthropic Claude 4 family2026-05-28
LMArena (Chatbot Arena)Human preference#2/ 1392LMArena leaderboard2026-05-25
GDPvalAgentic#2/ 6100Artificial Analysis — GDPval-AA v22026-06-17
HELM SafetySafety#2/ 992Stanford HELM leaderboard2026-05-28
LiveBenchReasoning#2/ 2171LiveBench leaderboard2026-04-15
MCP AtlasTool use#2/ 1079MCP Atlas leaderboard2026-06-09
SWE-bench VerifiedCoding#2/ 2389SWE-bench official leaderboard2026-05-30
BFCLTool use#3/ 2087Berkeley Function Calling Leaderboard V42026-04-15
DeepSWECoding#3/ 659DeepSWE leaderboard v1.12026-06-20
Humanity's Last ExamReasoning#3/ 6546Artificial Analysis2026-07-21
HumanEvalCoding#3/ 1596Anthropic Claude 4 family2026-05-28
MMMUMultimodal#3/ 871MMMU leaderboard2026-05-28
Artificial Analysis Intelligence IndexReasoning#4/ 6756Artificial Analysis2026-07-21
Vals IndexAgentic#5/ 2070Vals AI — Vals Index2026-06-04
MATHMath#6/ 2088Anthropic Claude 4 family2026-05-28
MMLU-ProKnowledge#9/ 2182MMLU-Pro HuggingFace leaderboard2026-05-28
AA time to first tokenPerformance#13/ 69Artificial Analysis2026-07-21
SciCodeCoding#14/ 6554Artificial Analysis2026-07-21
Terminal-BenchAgentic#17/ 6975Terminal-Bench 2.0 leaderboard2026-05-20
τ³-Bench BankingAgentic#30/ 6428Artificial Analysis2026-07-21
IFBenchReasoning#41/ 5762Artificial Analysis2026-07-21
AA output speedPerformance#45/ 69Artificial Analysis2026-07-21
AA-LCRLong context#48/ 6458Artificial Analysis — AA-LCR2026-06-17
GPQA DiamondReasoning#51/ 6980GPQA Diamond paper leaderboard2026-05-15
Family context

The three highest-scoring carded models in each capability family. Where Claude Opus 4.8 shows up, it's highlighted.

KnowledgeClaude Opus 4.8 · 91
  1. 1
    91
  2. 2
    91
  3. 3
    91
ReasoningClaude Opus 4.8 · 80
  1. 1
    95
  2. 2
    94
  3. 3
    94
MathClaude Opus 4.8 · 94
  1. 1
    99
  2. 2
    97
  3. 3
    96
CodingClaude Opus 4.8 · 96
  1. 1
    96
  2. 2
    96
  3. 3
    96
AgenticClaude Opus 4.8 · 100
  1. 1
    100
  2. 2
    100
  3. 3
    100
MultimodalClaude Opus 4.8 · 71
  1. 1
    72
  2. 2
    72
  3. 3
    71
Long contextClaude Opus 4.8 · 58
  1. 1
    75
  2. 2
    75
  3. 3
    74
Tool useClaude Opus 4.8 · 87
  1. 1
    89
  2. 2
    87
  3. 3
    87
SafetyClaude Opus 4.8 · 92
  1. 1
    92
  2. 2
    92
  3. 3
    91
Human preferenceClaude Opus 4.8 · 92
  1. 1
    95
  2. 2
    92
  3. 3
    91
Newest receipts

Benchmark rows added to the public ledger for Claude Opus 4.8 in the last 120 days. Older rows live in the full table below.

26 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Knowledge91
Reasoning80
Math94
Coding96
Agentic100
Multimodal71
Long context58
Tool use87
Safety92
Human preference92
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMiniMaxxAIMetaZhipuXiaomiMeituanCohere

Full benchmark ledger

Every catalog benchmark for Claude Opus 4.8. Scores link to the original source; gaps mean no public row exists yet.

26 sourced rows

26 of 32 catalog benchmarks have a sourced row for Claude Opus 4.8.

Knowledge

Reasoning

Math

Coding

Agentic

Multimodal

Long context

Tool use

Performance

Safety

Human preference

Changelog
  1. Released Claude Opus 4.8 — top of Vals (70.17%) and AA (61.4).
  2. Claude Opus 4.7 retired for general availability.
Sources