Models / Anthropic

Claude Sonnet 4.6

Claude Sonnet · Released 2026-02-17

Anthropic's February 2026 mid-tier model. Vals Index 60.30% (#3 on the public leaderboard — beats GPT 5.4 Mini and Gemini 3 Flash). SWE-bench Verified 79.6%, ARC-AGI-2 58.3%. 1M context, $3 / $15 per 1M tokens — 5× cheaper than Opus 4.8 on input.

The default for production work that needs Anthropic's voice and tool-use discipline without the Opus 4.8 price tag. SWE-bench Verified 79.6% makes it strong enough for most coding agents.

#8 of 21 on Vals Index · current snapshot
60 Vals Index · 24 benchmark rows · 11 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite66
Context1M
Input / 1M$3.0
Output / 1M$15
Knowledge cutoff2025-12-01
Output speed60 t/s
TTFT (default API)1.18s
Statusflagship
KnowledgeReasoningMathCodingAgenticMultimodalLong contextTool useSafetyHuman preference
Knowledge
91
Reasoning
77
Math
84
Coding
94
Agentic
80
Multimodal
66
Long context
58
Tool use
89
Safety
89
Human preference
79

81average across 10 tested families

Editor's note
Sonnet 4.6 is the workhorse. It leads Anthropic's mid-tier and sits at #3 on Vals — between Opus 4.8 and the GPT 5.4 family. For production agents where Opus 4.8 is overkill, this is the Anthropic pick.
Benchmark placements

Where Claude Sonnet 4.6 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
BFCLTool use#1/ 2089Berkeley Function Calling Leaderboard V42026-04-15
MCP AtlasTool use#4/ 1075MCP Atlas leaderboard2026-06-09
MMLUKnowledge#4/ 2291Anthropic Claude 3.7/4.6 model card2026-02-17
HELM SafetySafety#5/ 989Stanford HELM leaderboard2026-02-17
LiveBenchReasoning#5/ 2167LiveBench leaderboard2026-04-15
MMMUMultimodal#5/ 866MMMU leaderboard2026-02-17
HumanEvalCoding#7/ 1594Anthropic Claude 3.7/4.6 model card2026-01-01
ARC-AGIReasoning#8/ 1458ARC Prize semi-private evaluation2026-02-17
LongBench v2Long context#8/ 1953LongBench v2 leaderboard2026-04-01
Vals IndexAgentic#8/ 2060Vals AI — Vals Index2026-06-04
LMArena (Chatbot Arena)Human preference#10/ 1379LMArena leaderboard2026-05-25
MATHMath#11/ 2084Anthropic Claude 3.7/4.6 model card2026-02-17
τ³-Bench BankingAgentic#12/ 6480Artificial Analysis2026-07-21
MMLU-ProKnowledge#13/ 2178MMLU-Pro HuggingFace leaderboard2026-02-17
SWE-bench VerifiedCoding#13/ 2380SWE-bench official leaderboard2026-05-30
Terminal-BenchAgentic#24/ 6968Terminal-Bench 2.0 leaderboard2026-06-09
SciCodeCoding#37/ 6547Artificial Analysis2026-07-21
AA time to first tokenPerformance#39/ 69Artificial Analysis2026-07-21
Artificial Analysis Intelligence IndexReasoning#43/ 6736Artificial Analysis2026-07-21
AA output speedPerformance#48/ 69Artificial Analysis2026-07-21
AA-LCRLong context#51/ 6458Artificial Analysis2026-07-21
GPQA DiamondReasoning#54/ 6977Artificial Analysis GPQA Diamond evaluation2026-06-09
Humanity's Last ExamReasoning#54/ 6513Artificial Analysis2026-07-21
IFBenchReasoning#54/ 5741Artificial Analysis2026-07-21
Family context

The three highest-scoring carded models in each capability family. Where Claude Sonnet 4.6 shows up, it's highlighted.

KnowledgeClaude Sonnet 4.6 · 91
  1. 1
    91
  2. 2
    91
  3. 3
    91
ReasoningClaude Sonnet 4.6 · 77
  1. 1
    95
  2. 2
    94
  3. 3
    94
MathClaude Sonnet 4.6 · 84
  1. 1
    99
  2. 2
    97
  3. 3
    96
CodingClaude Sonnet 4.6 · 94
  1. 1
    96
  2. 2
    96
  3. 3
    96
AgenticClaude Sonnet 4.6 · 80
  1. 1
    100
  2. 2
    100
  3. 3
    100
MultimodalClaude Sonnet 4.6 · 66
  1. 1
    72
  2. 2
    72
  3. 3
    71
Long contextClaude Sonnet 4.6 · 58
  1. 1
    75
  2. 2
    75
  3. 3
    74
Tool useClaude Sonnet 4.6 · 89
  1. 1
    89
  2. 2
    87
  3. 3
    87
SafetyClaude Sonnet 4.6 · 89
  1. 1
    92
  2. 2
    92
  3. 3
    91
Human preferenceClaude Sonnet 4.6 · 79
  1. 1
    95
  2. 2
    92
  3. 3
    91
Newest receipts

Benchmark rows added to the public ledger for Claude Sonnet 4.6 in the last 120 days. Older rows live in the full table below.

17 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Knowledge91
Reasoning77
Math84
Coding94
Agentic80
Multimodal66
Long context58
Tool use89
Safety89
Human preference79
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMiniMaxxAIMetaZhipuXiaomiMeituanCohere

Full benchmark ledger

Every catalog benchmark for Claude Sonnet 4.6. Scores link to the original source; gaps mean no public row exists yet.

24 sourced rows

24 of 32 catalog benchmarks have a sourced row for Claude Sonnet 4.6.

Knowledge

Reasoning

Math

Coding

Agentic

Multimodal

Long context

Tool use

Performance

Safety

Human preference

Changelog
  1. Released Claude Sonnet 4.6 — Vals #3 at 60.30%, 1M context, $3 / $15.
Sources