Models / Moonshot

Kimi K3

Kimi K3 · Released 2026-07-16

Moonshot AI's 2.8T-parameter multimodal reasoning model for long-horizon coding and knowledge work, with a 1M-token context window, public API access, and open weights on Hugging Face.

Open-frontier option for large-repository coding, tool use, visual reasoning, and context-heavy research workflows — via API or self-hosted weights.

#1 of 24 on Vals Index · current snapshot
75 Vals Index · 14 benchmark rows · 3 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite58
Context1M
Input / 1M$3.0
Output / 1M$15
Knowledge cutoffunpublished
Output speed40 t/s
TTFT (default API)2.39s
Statussolid

Family profile

Best published score in each covered benchmark family.

75/100 avg
Knowledge
48
Reasoning
94
Coding
59
Agentic
85
Multimodal
81
Long context
83

6 tested benchmark families

Benchmark placements

Where Kimi K3 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
Vals IndexAgentic#1/ 2375Vals AI — Vals Index2026-07-16
AA-LCRLong context#2/ 17983Artificial Analysis2026-09-01
SciCodeCoding#3/ 18059Artificial Analysis2026-09-01
Artificial Analysis Intelligence IndexReasoning#4/ 18760Artificial Analysis2026-09-01
τ³-Bench BankingAgentic#7/ 10646Artificial Analysis2026-09-01
Terminal-Bench-ScienceAgentic#7/ 97Terminal-Bench-Science 0.1 announcement2026-08-27
GPQA DiamondReasoning#8/ 18594Artificial Analysis2026-09-01
MMMU-ProMultimodal#8/ 881Artificial Analysis — MMMU-Pro2026-09-01
CritPtReasoning#9/ 1223Artificial Analysis — CritPt2026-09-01
Humanity's Last ExamReasoning#10/ 18047Artificial Analysis2026-09-01
Terminal-BenchAgentic#10/ 18485Artificial Analysis2026-09-01
AA-OmniscienceKnowledge#11/ 1248Artificial Analysis — AA-Omniscience2026-09-01
AA time to first tokenPerformance#26/ 186Artificial Analysis2026-09-01
AA output speedPerformance#81/ 187Artificial Analysis2026-09-01
Family context

The three highest-scoring models with pages in each capability family. Where Kimi K3 shows up, it's highlighted.

Newest receipts

Benchmark rows added to the public ledger for Kimi K3 in the last 120 days. Older rows live in the full table below.

14 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Knowledge48
Reasoning94
Coding59
Agentic85
Multimodal81
Long context83
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover any dot for name, score, and input price. Keyboard: tab through the top twelve, or use the ranking below.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMetaMiniMaxxAIZhipuXiaomiOtherMeituanCohereNVIDIA

Full benchmark ledger

Every catalog benchmark for Kimi K3. Scores link to the original source; gaps mean no public row exists yet.

14 sourced rows

14 of 37 catalog benchmarks have a sourced row for Kimi K3.

Knowledge

Reasoning

Coding

Agentic

Multimodal

Long context

Performance

Sources