Anthropic's May 2026 flagship. Tops the Vals Index (70.17%) and the Artificial Analysis Intelligence Index (61.4). SWE-bench Verified 88.6%, Terminal-Bench 2.0 74.6%, 1M-token context, $15 / $75 per 1M tokens.
Front-door model for hard reasoning, long documents, agentic coding, and tool-heavy research. The model we reach for first when a task is too hard for Sonnet but small enough to fit in one prompt.
#4 of 24 on Vals Index · current snapshot
70 Vals Index · 30 benchmark rows · 12 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Best published score in each covered benchmark family.
86/100 avg
050100
Knowledge
91
Reasoning
80
Math
94
Coding
96
Agentic
100
Multimodal
71
Long context
58
Tool use
87
Safety
92
Human preference
92
10 tested benchmark families
Editor's note
Opus 4.8 is the model we reach for first when a task is hard enough to make a Sonnet hallucinate but small enough to fit in one prompt. The 1M context window is real, not marketing — long-document citation fidelity and BFCL scores back it up. The single-number rank (Vals, AA) is the highest we've measured.
Benchmark placements
Per-benchmark positions
Where Claude Opus 4.8 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.