Modelle / Anthropic

Claude Sonnet 4.6

Claude Sonnet · Release 2026-02-17

Anthropics Mittelklassemodell Februar 2026. Vals-Index 60,30 % (#3 auf dem öffentlichen Leaderboard — schlägt GPT 5.4 Mini und Gemini 3 Flash). SWE-bench Verified 79,6 %, ARC-AGI-2 58,3 %. 1M-Kontext, $3 / $15 pro 1M Tokens — 5× günstiger als Opus 4.8 beim Input.

Standard für Produktionsarbeit, die Anthropics Stimme und Tool-Use-Disziplin braucht, ohne den Opus-4.8-Preis. SWE-bench Verified 79,6 % machen es stark genug für die meisten Coding-Agents.

#8 von 21 im Vals Index · aktueller Snapshot
60 Vals Index · 24 Benchmark-Zeilen · 11 QuellenPosition ist benchmark-spezifisch — kein Cross-Family- oder Cross-Source-Ranking.
Composite66
Kontext1M
Input / 1M$3.0
Output / 1M$15
Wissensstichtag2025-12-01
Output-Speed60 t/s
TTFT (Default-API)1.18s
Statusflagship
KnowledgeReasoningMathCodingAgenticMultimodalLong contextTool useSafetyHuman preference
Knowledge
91
Reasoning
77
Math
84
Coding
94
Agentic
80
Multimodal
66
Long context
58
Tool use
89
Safety
89
Human preference
79

81average across 10 tested families

Redaktionsnotiz
Sonnet 4.6 ist das Arbeitstier. Es führt Anthropics Mittelklasse an und liegt auf #3 bei Vals — zwischen Opus 4.8 und der GPT-5.4-Familie. Für Produktions-Agents, bei denen Opus 4.8 zu viel ist, ist dies die Anthropic-Wahl.
Benchmark-Platzierungen

Wo Claude Sonnet 4.6 auf jeder öffentlichen Benchmark-Quelle mit Zeile steht. Rang zählt jedes Modell mit neuester Zeile im selben Test — kein universeller Qualitätsscore.

BenchmarkFamilyRangScoreQuelleDatum
BFCLTool use#1/ 2089Berkeley Function Calling Leaderboard V42026-04-15
MCP AtlasTool use#4/ 1075MCP Atlas leaderboard2026-06-09
MMLUKnowledge#4/ 2291Anthropic Claude 3.7/4.6 model card2026-02-17
HELM SafetySafety#5/ 989Stanford HELM leaderboard2026-02-17
LiveBenchReasoning#5/ 2167LiveBench leaderboard2026-04-15
MMMUMultimodal#5/ 866MMMU leaderboard2026-02-17
HumanEvalCoding#7/ 1594Anthropic Claude 3.7/4.6 model card2026-01-01
ARC-AGIReasoning#8/ 1458ARC Prize semi-private evaluation2026-02-17
LongBench v2Long context#8/ 1953LongBench v2 leaderboard2026-04-01
Vals IndexAgentic#8/ 2060Vals AI — Vals Index2026-06-04
LMArena (Chatbot Arena)Human preference#10/ 1379LMArena leaderboard2026-05-25
MATHMath#11/ 2084Anthropic Claude 3.7/4.6 model card2026-02-17
τ³-Bench BankingAgentic#12/ 6480Artificial Analysis2026-07-21
MMLU-ProKnowledge#13/ 2178MMLU-Pro HuggingFace leaderboard2026-02-17
SWE-bench VerifiedCoding#13/ 2380SWE-bench official leaderboard2026-05-30
Terminal-BenchAgentic#24/ 6968Terminal-Bench 2.0 leaderboard2026-06-09
SciCodeCoding#37/ 6547Artificial Analysis2026-07-21
AA time to first tokenPerformance#39/ 69Artificial Analysis2026-07-21
Artificial Analysis Intelligence IndexReasoning#43/ 6736Artificial Analysis2026-07-21
AA output speedPerformance#48/ 69Artificial Analysis2026-07-21
AA-LCRLong context#51/ 6458Artificial Analysis2026-07-21
GPQA DiamondReasoning#54/ 6977Artificial Analysis GPQA Diamond evaluation2026-06-09
Humanity's Last ExamReasoning#54/ 6513Artificial Analysis2026-07-21
IFBenchReasoning#54/ 5741Artificial Analysis2026-07-21
Family-Kontext

Die drei höchstscorierenden kartierten Modelle pro Capability-Family. Wo Claude Sonnet 4.6 auftaucht, ist es hervorgehoben.

KnowledgeClaude Sonnet 4.6 · 91
  1. 1
    91
  2. 2
    91
  3. 3
    91
ReasoningClaude Sonnet 4.6 · 77
  1. 1
    95
  2. 2
    94
  3. 3
    94
MathClaude Sonnet 4.6 · 84
  1. 1
    99
  2. 2
    97
  3. 3
    96
CodingClaude Sonnet 4.6 · 94
  1. 1
    96
  2. 2
    96
  3. 3
    96
AgenticClaude Sonnet 4.6 · 80
  1. 1
    100
  2. 2
    100
  3. 3
    100
MultimodalClaude Sonnet 4.6 · 66
  1. 1
    72
  2. 2
    72
  3. 3
    71
Long contextClaude Sonnet 4.6 · 58
  1. 1
    75
  2. 2
    75
  3. 3
    74
Tool useClaude Sonnet 4.6 · 89
  1. 1
    89
  2. 2
    87
  3. 3
    87
SafetyClaude Sonnet 4.6 · 89
  1. 1
    92
  2. 2
    92
  3. 3
    91
Human preferenceClaude Sonnet 4.6 · 79
  1. 1
    95
  2. 2
    92
  3. 3
    91
Neueste Belege

Benchmark-Zeilen im öffentlichen Ledger für Claude Sonnet 4.6 in den letzten 120 Tagen. Ältere Zeilen stehen in der vollen Tabelle unten.

17 neu
Family-Abdeckung

Ein hoher Composite, der eine schwache Family versteckt, ist eine Falle. Diese Balken zeigen Families ohne öffentlichen Test — und wo das Modell führt.

Knowledge91
Reasoning77
Math84
Coding94
Agentic80
Multimodal66
Long context58
Tool use89
Safety89
Human preference79
Preis vs. Performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMiniMaxxAIMetaZhipuXiaomiMeituanCohere

Volles Benchmark-Ledger

Jeder Katalog-Benchmark für Claude Sonnet 4.6. Scores verlinken zur Originalquelle; Lücken heißen: noch keine öffentliche Zeile.

24 Quell-Zeilen

24 of 32 catalog benchmarks have a sourced row for Claude Sonnet 4.6.

Knowledge

Reasoning

Math

Coding

Agentic

Multimodal

Long context

Tool use

Performance

Safety

Human preference

Changelog
  1. Claude Sonnet 4.6 veröffentlicht — Vals #3 bei 60,30 %, 1M-Kontext, $3 / $15.
Quellen