Models / Moonshot

Kimi K2.5

Kimi · Released 2026-01-27

Moonshot's January 2026 multimodal open-weights model. 262K context, $0.60/$3 per 1M tokens. Agent Swarm up to 100 sub-agents. Predecessor to K2.6.

Open-weights Moonshot pick when K2.6 isn't available or you need the K2.5 weight checkpoint.

#85 of 168 on AA Index · current snapshot
36 AA Index · 13 benchmark rows · 4 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite54
Context262K
Input / 1M$0.60
Output / 1M$3.0
Knowledge cutoff2025-10-01
Statussolid

Family profile

Best published score in each covered benchmark family.

68/100 avg
Knowledge
86
Reasoning
88
Coding
49
Agentic
46
Long context
73

5 tested benchmark families

Benchmark placements

Where Kimi K2.5 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
LongBench v2Long context#15/ 1948LongBench v2 leaderboard2026-04-01
LiveBenchReasoning#16/ 2161LiveBench leaderboard2026-04-15
MMLUKnowledge#18/ 2286Moonshot Kimi K2.52026-01-27
SciCodeCoding#51/ 18049Artificial Analysis2026-09-01
IFBenchReasoning#54/ 12770Artificial Analysis2026-09-01
GPQA DiamondReasoning#64/ 18588Artificial Analysis2026-09-01
AA-LCRLong context#68/ 17973Artificial Analysis2026-09-01
Humanity's Last ExamReasoning#70/ 18031Artificial Analysis2026-09-01
τ³-Bench BankingAgentic#75/ 10614Artificial Analysis2026-09-01
Terminal-BenchAgentic#97/ 18446Artificial Analysis2026-09-01
Artificial Analysis Intelligence IndexReasoning#101/ 18736Artificial Analysis2026-09-01
AA time to first tokenPerformance#158/ 186Artificial Analysis2026-09-01
AA output speedPerformance#159/ 187Artificial Analysis2026-09-01
Family context

The three highest-scoring models with pages in each capability family. Where Kimi K2.5 shows up, it's highlighted.

Newest receipts

Benchmark rows added to the public ledger for Kimi K2.5 in the last 120 days. Older rows live in the full table below.

10 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Knowledge86
Reasoning88
Coding49
Agentic46
Long context73
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover any dot for name, score, and input price. Keyboard: tab through the top twelve, or use the ranking below.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMetaMiniMaxxAIZhipuXiaomiOtherMeituanCohereNVIDIA

Full benchmark ledger

Every catalog benchmark for Kimi K2.5. Scores link to the original source; gaps mean no public row exists yet.

13 sourced rows

13 of 37 catalog benchmarks have a sourced row for Kimi K2.5.

Knowledge

Reasoning

Coding

Agentic

Long context

Performance

Sources