Models / Moonshot

Kimi K2.6

Kimi · Released 2026-04-20

Moonshot's April 2026 open-weights flagship. Vals Index 55.55% (#5 on the public leaderboard), SWE-bench Verified 80.2%, Terminal-Bench 2.0 66.7%. Leads the open-source field on GPQA Diamond at 90.5%. 256K context, $0.60 / $2.50 per 1M tokens.

Default open-weights model from a Chinese provider for English-language work. The strongest open-source model on GPQA Diamond (90.5%) and Vals (55.55%) as of June 2026.

#13 of 24 on Vals Index · current snapshot
56 Vals Index · 23 benchmark rows · 10 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite67
Context256K
Input / 1M$0.60
Output / 1M$2.5
Knowledge cutoff2026-02-15
Statusflagship

Family profile

Best published score in each covered benchmark family.

83/100 avg
Knowledge
89
Reasoning
91
Math
85
Coding
94
Agentic
67
Long context
77
Tool use
85
Human preference
78

8 tested benchmark families

Editor's note
Kimi K2.6 is the strongest open-weights model we have on the public ledger, and the only open-source model on the Vals top 10. If you need to self-host a frontier-tier model and English-language support matters, this is the one.
Benchmark placements

Where Kimi K2.6 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
BFCLTool use#4/ 2085Berkeley Function Calling Leaderboard V42026-04-15
HumanEvalCoding#8/ 1594CodeSOTA HumanEval leaderboard2026-04-20
MCP AtlasTool use#9/ 1070MCP Atlas leaderboard2026-06-09
MMLUKnowledge#9/ 2289Moonshot Kimi K2.62026-04-20
SWE-bench VerifiedCoding#9/ 2380SWE-bench official leaderboard2026-05-30
LMArena (Chatbot Arena)Human preference#11/ 1378LMArena leaderboard2026-05-25
FrontierMathMath#11/ 1539Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15
LiveBenchReasoning#12/ 2162LiveBench leaderboard2026-04-15
LongBench v2Long context#13/ 1949LongBench v2 leaderboard2026-04-01
Vals IndexAgentic#13/ 2356Vals AI — Vals Index2026-06-04
MATHMath#18/ 2583Moonshot Kimi K2.62026-04-20
IFBenchReasoning#19/ 12776Artificial Analysis2026-09-01
SciCodeCoding#24/ 18054Artificial Analysis2026-09-01
AIMEMath#25/ 3785Moonshot Kimi K2.62026-04-20
AA-LCRLong context#35/ 17977Artificial Analysis2026-09-01
MMLU-ProKnowledge#35/ 3975MMLU-Pro HuggingFace leaderboard2026-04-20
GPQA DiamondReasoning#36/ 18591GPQA Diamond paper leaderboard2026-05-15
Humanity's Last ExamReasoning#45/ 18038Artificial Analysis2026-09-01
Artificial Analysis Intelligence IndexReasoning#48/ 18745Artificial Analysis2026-09-01
τ³-Bench BankingAgentic#49/ 10623Artificial Analysis2026-09-01
Terminal-BenchAgentic#53/ 18467Terminal-Bench 2.0 leaderboard2026-05-20
AA time to first tokenPerformance#159/ 186Artificial Analysis2026-09-01
AA output speedPerformance#160/ 187Artificial Analysis2026-09-01
Family context

The three highest-scoring models with pages in each capability family. Where Kimi K2.6 shows up, it's highlighted.

KnowledgeKimi K2.6 · 89
  1. 1
    91
  2. 2
    91
  3. 3
    91
ReasoningKimi K2.6 · 91
  1. 1
    95
  2. 2
    95
  3. 3
    95
MathKimi K2.6 · 85
  1. 1
    99
  2. 2
    99
  3. 3
    99
CodingKimi K2.6 · 94
  1. 1
    96
  2. 2
    96
  3. 3
    96
AgenticKimi K2.6 · 67
  1. 1
    100
  2. 2
    100
  3. 3
    100
Long contextKimi K2.6 · 77
  1. 1
    83
  2. 2
    83
  3. 3
    81
Tool useKimi K2.6 · 85
  1. 1
    89
  2. 2
    87
  3. 3
    87
Human preferenceKimi K2.6 · 78
  1. 1
    95
  2. 2
    92
  3. 3
    91
Newest receipts

Benchmark rows added to the public ledger for Kimi K2.6 in the last 120 days. Older rows live in the full table below.

15 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Knowledge89
Reasoning91
Math85
Coding94
Agentic67
Long context77
Tool use85
Human preference78
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover any dot for name, score, and input price. Keyboard: tab through the top twelve, or use the ranking below.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMetaMiniMaxxAIZhipuXiaomiOtherMeituanCohereNVIDIA

Full benchmark ledger

Every catalog benchmark for Kimi K2.6. Scores link to the original source; gaps mean no public row exists yet.

23 sourced rows

23 of 37 catalog benchmarks have a sourced row for Kimi K2.6.

Knowledge

Reasoning

Math

Coding

Agentic

Long context

Tool use

Performance

Human preference

Changelog
  1. Released Kimi K2.6 — Vals #5 at 55.55%, GPQA Diamond 90.5% (open-source leader).
Sources