90.6%
MMLU
Modelle / Anthropic
Claude Sonnet · Release 2026-02-17
Anthropics Mittelklassemodell Februar 2026. Vals-Index 60,30 % (#3 auf dem öffentlichen Leaderboard — schlägt GPT 5.4 Mini und Gemini 3 Flash). SWE-bench Verified 79,6 %, ARC-AGI-2 58,3 %. 1M-Kontext, $3 / $15 pro 1M Tokens — 5× günstiger als Opus 4.8 beim Input.
Standard für Produktionsarbeit, die Anthropics Stimme und Tool-Use-Disziplin braucht, ohne den Opus-4.8-Preis. SWE-bench Verified 79,6 % machen es stark genug für die meisten Coding-Agents.
81average across 10 tested families
Wo Claude Sonnet 4.6 auf jeder öffentlichen Benchmark-Quelle mit Zeile steht. Rang zählt jedes Modell mit neuester Zeile im selben Test — kein universeller Qualitätsscore.
Die drei höchstscorierenden kartierten Modelle pro Capability-Family. Wo Claude Sonnet 4.6 auftaucht, ist es hervorgehoben.
Benchmark-Zeilen im öffentlichen Ledger für Claude Sonnet 4.6 in den letzten 120 Tagen. Ältere Zeilen stehen in der vollen Tabelle unten.
| Benchmark | Family | Score | Quelle | Tage her |
|---|---|---|---|---|
| AA-LCR | Long context | 57.67% | Artificial Analysis | 2 |
| AA output speed | Performance | 60 t/s | Artificial Analysis | 2 |
| AA time to first token | Performance | 1.18s | Artificial Analysis | 2 |
| Artificial Analysis Intelligence Index | Reasoning | 35.9% | Artificial Analysis | 2 |
| Humanity's Last Exam | Reasoning | 13.2% | Artificial Analysis | 2 |
| IFBench | Reasoning | 41.16% | Artificial Analysis | 2 |
| SciCode | Coding | 46.9% | Artificial Analysis | 2 |
| τ³-Bench Banking | Agentic | 79.53% | Artificial Analysis | 2 |
| Terminal-Bench | Agentic | 68.4% | Terminal-Bench 2.0 leaderboard | 44 |
| GPQA Diamond | Reasoning | 77.2% | Artificial Analysis GPQA Diamond evaluation | 44 |
| MCP Atlas | Tool use | 75.4% | MCP Atlas leaderboard | 44 |
| Vals Index | Agentic | 60.3% | Vals AI — Vals Index | 49 |
| SWE-bench Verified | Coding | 79.6% | SWE-bench official leaderboard | 54 |
| LMArena (Chatbot Arena) | Human preference | 1396 | LMArena leaderboard | 59 |
| BFCL | Tool use | 88.7% | Berkeley Function Calling Leaderboard V4 | 99 |
| LiveBench | Reasoning | 67.2% | LiveBench leaderboard | 99 |
| LongBench v2 | Long context | 52.6% | LongBench v2 leaderboard | 113 |
Ein hoher Composite, der eine schwache Family versteckt, ist eine Falle. Diese Balken zeigen Families ohne öffentlichen Test — und wo das Modell führt.
Hover or tab any dot for name, score, and input price.
Jeder Katalog-Benchmark für Claude Sonnet 4.6. Scores verlinken zur Originalquelle; Lücken heißen: noch keine öffentliche Zeile.
24 of 32 catalog benchmarks have a sourced row for Claude Sonnet 4.6.
58.3%
ARC-AGI-2 semi-private
35.9%
AA Intelligence Index v4.1
77.2%
GPQA Diamond
13.2%
HLE (AA run)
41.16%
IFBench (AA run)
67.2%
LiveBench
84.2%
MATH
94.1% pass@1
HumanEval
46.9%
SciCode (AA run)
79.6%
SWE-bench Verified
68.4%
Terminal-Bench 2.0 (audited harness)
60.3%
Vals Index
79.53%
τ²-Bench Banking (AA run; maps to τ³ explainer)
66.2%
MMMU
57.67%
AA-LCR (AA run)
52.6%
LongBench v2
60 t/s
Median output tokens/s (1k prompt, default provider)
1.18s
Median time to first token (1k prompt, default provider)
89.2%
HELM Safety
1396
LMArena Elo (Style Controlled)