Models / Moonshot

Kimi K2.6

Kimi · Released 2026-04-20

Moonshot's April 2026 open-weights flagship. Vals Index 55.55% (#5 on the public leaderboard), SWE-bench Verified 80.2%, Terminal-Bench 2.0 66.7%. Leads the open-source field on GPQA Diamond at 90.5%. 256K context, $0.60 / $2.50 per 1M tokens.

Default open-weights model from a Chinese provider for English-language work. The strongest open-source model on GPQA Diamond (90.5%) and Vals (55.55%) as of June 2026.

#10 of 21 on Vals Index · current snapshot
56 Vals Index · 22 benchmark rows · 10 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite67
Context256K
Input / 1M$0.60
Output / 1M$2.5
Knowledge cutoff2026-02-15
Output speed56 t/s
TTFT (default API)1.08s
Statusflagship
KnowledgeReasoningMathCodingAgenticLong contextTool useHuman preference
Knowledge
89
Reasoning
91
Math
85
Coding
94
Agentic
67
Long context
70
Tool use
85
Human preference
78

82average across 8 tested families

Editor's note
Kimi K2.6 is the strongest open-weights model we have on the public ledger, and the only open-source model on the Vals top 10. If you need to self-host a frontier-tier model and English-language support matters, this is the one.
Benchmark placements

Where Kimi K2.6 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
BFCLTool use#4/ 2085Berkeley Function Calling Leaderboard V42026-04-15
HumanEvalCoding#8/ 1594CodeSOTA HumanEval leaderboard2026-04-20
MCP AtlasTool use#9/ 1070MCP Atlas leaderboard2026-06-09
MMLUKnowledge#9/ 2289Moonshot Kimi K2.62026-04-20
SWE-bench VerifiedCoding#9/ 2380SWE-bench official leaderboard2026-05-30
Vals IndexAgentic#10/ 2056Vals AI — Vals Index2026-06-04
AIMEMath#11/ 1985Moonshot Kimi K2.62026-04-20
LMArena (Chatbot Arena)Human preference#11/ 1378LMArena leaderboard2026-05-25
LiveBenchReasoning#12/ 2162LiveBench leaderboard2026-04-15
IFBenchReasoning#13/ 5776Artificial Analysis2026-07-21
LongBench v2Long context#13/ 1949LongBench v2 leaderboard2026-04-01
MATHMath#13/ 2083Moonshot Kimi K2.62026-04-20
SciCodeCoding#16/ 6554Artificial Analysis2026-07-21
MMLU-ProKnowledge#17/ 2175MMLU-Pro HuggingFace leaderboard2026-04-20
GPQA DiamondReasoning#21/ 6991GPQA Diamond paper leaderboard2026-05-15
AA-LCRLong context#23/ 6470Artificial Analysis2026-07-21
Artificial Analysis Intelligence IndexReasoning#24/ 6744Artificial Analysis2026-07-21
Humanity's Last ExamReasoning#26/ 6536Artificial Analysis2026-07-21
Terminal-BenchAgentic#27/ 6967Terminal-Bench 2.0 leaderboard2026-05-20
τ³-Bench BankingAgentic#40/ 6421Artificial Analysis2026-07-21
AA time to first tokenPerformance#42/ 69Artificial Analysis2026-07-21
AA output speedPerformance#52/ 69Artificial Analysis2026-07-21
Family context

The three highest-scoring carded models in each capability family. Where Kimi K2.6 shows up, it's highlighted.

KnowledgeKimi K2.6 · 89
  1. 1
    91
  2. 2
    91
  3. 3
    91
ReasoningKimi K2.6 · 91
  1. 1
    95
  2. 2
    94
  3. 3
    94
MathKimi K2.6 · 85
  1. 1
    99
  2. 2
    97
  3. 3
    96
CodingKimi K2.6 · 94
  1. 1
    96
  2. 2
    96
  3. 3
    96
AgenticKimi K2.6 · 67
  1. 1
    100
  2. 2
    100
  3. 3
    100
Long contextKimi K2.6 · 70
  1. 1
    75
  2. 2
    75
  3. 3
    74
Tool useKimi K2.6 · 85
  1. 1
    89
  2. 2
    87
  3. 3
    87
Human preferenceKimi K2.6 · 78
  1. 1
    95
  2. 2
    92
  3. 3
    91
Newest receipts

Benchmark rows added to the public ledger for Kimi K2.6 in the last 120 days. Older rows live in the full table below.

22 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Knowledge89
Reasoning91
Math85
Coding94
Agentic67
Long context70
Tool use85
Human preference78
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMiniMaxxAIMetaZhipuXiaomiMeituanCohere

Full benchmark ledger

Every catalog benchmark for Kimi K2.6. Scores link to the original source; gaps mean no public row exists yet.

22 sourced rows

22 of 32 catalog benchmarks have a sourced row for Kimi K2.6.

Knowledge

Reasoning

Math

Coding

Agentic

Long context

Tool use

Performance

Human preference

Changelog
  1. Released Kimi K2.6 — Vals #5 at 55.55%, GPQA Diamond 90.5% (open-source leader).
Sources