VerdictPal · editorial desk · updated 5 Sep 2026VerdictPal
Models

Model-first evidence atlas.

We score models, not vendors. Below: every tooled model the public benchmark ledger has evidence for, sorted by intelligence receipts (Vals Index first, then Artificial Analysis). Every row carries a date and a URL, not a universal verdict.
Current Vals evidence anchor · Vals Index
Kimi K3Moonshot · Kimi K3 · 75 on the Vals intelligence index
216
Models with pages
21
Providers
171
Priced in USD/1M
8
In score ledger only

Vals Index rows appear first. Models without a Vals row fall back to Artificial Analysis. Source badges show whether a row is single-source or corroborated — not a cross-benchmark winner.

  1. 1
    Kimi K3Moonshot · Kimi K3 · Vals + AA agree
    75Vals
  2. 2
    Claude Mythos PreviewAnthropic · Claude Opus · Vals only
    73Vals
  3. 3
    GPT-5.6 SolOpenAI · GPT-5.6 · Vals + AA agree
    73Vals
  4. 4
    Claude Opus 4.8Anthropic · Claude Opus · Vals + AA agree
    70Vals
  5. 5
    Claude Fable 5.1Anthropic · Claude Fable · Vals + AA agree
    69Vals
  6. 6
    GPT 5.5OpenAI · GPT 5 · Vals + AA agree
    68Vals
  7. 7
    Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic · Claude Opus · Vals + AA agree
    67Vals
  8. 8
    GPT-6 AstraOpenAI · GPT-6 · Vals + AA agree
    67Vals
  9. 9
    Claude Opus 4.7Anthropic · Claude Opus · Vals + AA agree
    66Vals
  10. 10
    Claude Fable 5Anthropic · Claude Fable · Vals + AA agree
    66Vals
Top 10 of 192

Blended cost per million tokens (input + 3× output) on the X axis, log scale. Intelligence index on the Y axis — Vals when present, otherwise Artificial Analysis. Dots are sized by context window and connected by the Pareto frontier (best intelligence at each cost).

Pareto frontier · curated models159 scored · log scale
$0.10$1$10$1000255075100Blended $/1M (log)IntelligenceQwen 3.6 Plus
All 159 models in this chart
ModelIntelligenceBlended $/1MFrontier
Qwen 3.6 Plus79$1.63Yes
DeepSeek V4 Flash Max78$0.25Yes
Kimi K375$12.00
GPT-5.6 Sol73$16.00
Claude Mythos Preview73$243.75
Claude Opus 4.870$60.00
Claude Fable 5.169$40.00
GPT 5.568$40.63
Claude Opus 5 (Adaptive Reasoning, Max Effort)67$20.00
GPT-6 Astra67$40.00
Claude Opus 4.766$20.00
Claude Fable 566$40.00
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)63$20.00
Claude Opus 5 (Adaptive Reasoning, High Effort)62$20.00
GLM-5.360$3.65
Claude Sonnet 4.660$12.00
Claude Opus 5 (Adaptive Reasoning, Medium Effort)59$20.00
GLM-5.3-Flash58$0.41
Qwen3.8 2.4T A95B58$5.00
Qwen3.8 Max58$5.00
Muse Spark 1.257$3.50
GPT-5.6 Terra57$11.88
Qwen3.8-Flash-Next56$0.39
DeepSeek V4 Pro Max56$1.30
Kimi K2.656$2.02
Grok 4.556$5.00
Claude Sonnet 555$8.00
GPT-5.5 (high)55$23.75
GPT-5.4 Pro55$142.50
GPT-5.5 Pro54$142.50
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)53$3.30
Muse Spark 1.153$3.50
Muse Spark 1.353$3.50
GLM-5.253$3.65
Gemini 3.1 Pro53$17.50
Claude Opus 5 (Adaptive Reasoning, Low Effort)53$20.00
GPT 5.453$22.75
DeepSeek V4 Flash 0731 (Reasoning, Max Effort)52$1.10
DeepSeek V4 Flash Vision (Reasoning, Max Effort)52$1.10
Qwen3.8 27B52$2.38
Gemini 3.6 Flash (high)52$3.00
GPT-5.6 Luna52$4.75
GPT-5.4 Mini51$3.56
Grok 4.651$5.00
GPT-5.5 (medium)51$23.75
Agnes 2.5 Pro Beta49$0.25
Gemini 3.5 Flash49$0.97
Gemini 3 Flash Preview49$2.38
Qwen 3.7 Max48$1.00
Gemini 3.1 Pro Preview48$9.50
Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)48$12.00
Kimi K3 (low)48$12.00
Motif 347$0.00Yes
Gemini 3.8 Flash47$3.00
Qwen3.7 Max47$6.25
Gemini 3.5 Flash (medium)47$7.13
Grok 4.347$20.00
GPT-5.4 Nano46$0.99
GPT-5.3 Codex (xhigh)46$10.94
MiniMax M345$1.60
Gemini 3.7 Flash45$3.00
Claude Opus 4.6 (Adaptive Reasoning, Max Effort)45$20.00
GPT-5.5 (low)45$23.75
DeepSeek V4 Pro (Reasoning, High Effort)44$0.76
Apodex 1.144$2.33
Claude Opus 4.7 (Non-reasoning, High Effort)44$20.00
MiMo-V2.5-Pro43$0.76
Kimi K2.7 Code43$3.24
GPT-5.243$10.94
Hy342$0.00
Solar Pro 442$0.97
Nex-N2-Pro42$2.00
Inkling (xhigh)42$3.98
Claude Opus 4.5 (Reasoning)42$20.00
MiMo-V2-Pro41$0.00
Inkling Small41$0.97
MiniMax M2.741$1.21
Grok Build 0.1 061641$1.75
Qwen3.6 Plus41$2.38
GLM-5 (Reasoning)41$2.65
Qwen3.6 Max Preview41$6.17
GPT-5.2 Codex (xhigh)41$10.94
JT-4.1 Flash 236B A21B40$0.00
Agnes 2.5 Pro Alpha40$0.79
Claude Haiku 4.540$4.00
GPT-5.4 (low)40$11.88
GLM-5-Turbo39$0.00
DeepSeek V4 Flash (Reasoning, High Effort)39$0.25
Gemini 3 Flash Preview (Reasoning)39$2.38
GPT-5.2 (medium)39$10.94
Claude Opus 4.639$20.00
MiMo-V2.538$0.25
Qwen3.6 27B (Reasoning)38$2.85
MiMo-V2-Omni-032737$0.00
Solar Open2 250B37$0.00
Gemini 3.5 Flash-Lite37$1.95
Grok 4.3 (medium)37$2.19
Grok 4.20 0309 (Reasoning)37$5.00
GPT-5 Codex (high)37$7.81
Claude 4.5 Sonnet (Reasoning)37$12.00
MiMo-V2-Omni36$0.00
Grok 4.3 (low)36$2.19
Kimi K2.536$2.40
Gemini 3.5 Flash (minimal)36$7.13
GPT-5.1 Codex (high)36$7.81
Claude Opus 4.536$20.00
GPT-5.5 (Non-reasoning)36$23.75
A.X-K235$0.00
GLM 5V Turbo (Reasoning)35$0.00
MiniMax M2.535$0.97
Muse Glimmer35$1.09
Gemini 3.1 Flash Lite35$1.19
GLM-4.7 (Reasoning)35$1.80
Qwen3.5 27B (Reasoning)35$1.87
Kimi K2.6 (Non-reasoning)35$3.24
GLM-5.2 (Non-reasoning)35$3.53
GPT-5 (high)35$7.81
GPT-5 (medium)35$7.81
Claude Sonnet 4.6 (Non-reasoning, Low Effort)35$12.00
Claude 4.1 Opus (Reasoning)35$60.00
MiMo-V2-Flash (Feb 2026)34$0.00
Hy3-preview (Reasoning)34$0.19
KAT Coder Pro V234$0.97
Kimi K2 Thinking34$2.02
LongCat-2.034$2.40
Qwen3.5 397B A17B (Reasoning)34$2.85
Gemini 3 Pro Preview (low)34$9.50
Grok 434$12.00
GPT-5.5 Instant (May 2026)34$23.75
DeepSeek V4 Pro (Non-reasoning)32$0.76
Qwen3.6 35B A3B (Reasoning)32$1.18
Ring-2.6-1T32$1.95
K-EXAONE 2.0 080331$0.00
Qwen3.6 27B (Non-reasoning)31$2.85
Mistral Medium 3.530$6.00
JT-35B-Flash29$0.00
DeepSeek V4 Flash (Non-reasoning)29$0.25
GPT-5.5 Instant (June 2026)29$23.75
MiMo-V2.5-Pro (Non-reasoning)28$0.76
Hy3-preview (Non-reasoning)27$0.19
Ling-2.6-1T27$1.95
Qwen3.6 35B A3B (Non-reasoning)25$1.78
Grok 4.3 (Non-reasoning)25$2.19
Nemotron 3.5 Lightning24$0.16
Granite 4.2 30B24$0.53
Command A23$8.13
Grok 4.20 0309 v2 (Non-reasoning)22$5.00
EXAONE 4.5 33B21$0.00
Granite 4.2 8B20$0.20
JT-MINI19$0.00
G9v3-3B16$0.00
Mistral Large 316$1.25
Nemotron 3 Nano Omni 30B A3B Reasoning15$0.24
DiffusionGemma 26B A4B14$0.00
Granite 4.2 3B14$0.10
Ling 2.6 Flash14$0.25
MiniCPM5-1B (Non-reasoning)12$0.00
Granite 4.1 30B9$0.00
MiniCPM-V 4.6 1.3B4$0.00

The earlier cut, using input cost alone. Useful when your workload is input-heavy (RAG, document Q&A) rather than output-heavy.

0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover any dot for name, score, and input price. Keyboard: tab through the top twelve, or use the ranking below.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMetaMiniMaxxAIZhipuXiaomiOtherMeituanCohereNVIDIA

The highest intelligence receipt in each family slice. Use this as a sanity check that a single number isn't hiding a weak spot.

Kimi K375

Moonshot · Kimi K3

Family profile

Best published score in each covered benchmark family.

75/100 avg
Knowledge
48
Reasoning
94
Coding
59
Agentic
85
Multimodal
81
Long context
83

6 tested benchmark families

Claude Mythos Preview73

Anthropic · Claude Opus

Family profile

Best published score in each covered benchmark family.

91/100 avg
Reasoning
95
Coding
94
Agentic
84

3 tested benchmark families

GPT-5.6 Sol73

OpenAI · GPT-5.6

Family profile

Best published score in each covered benchmark family.

79/100 avg
Knowledge
59
Reasoning
94
Coding
73
Agentic
88
Multimodal
83
Long context
78

6 tested benchmark families

A model with no row in a family is a model that hasn't been publicly tested there. We surface that instead of pretending it's weak.

Open the benchmark atlas
Knowledge57models covered
Reasoning195models covered
Math51models covered
Coding189models covered
Agentic188models covered
Multimodal15models covered
Long context184models covered
Tool use20models covered
Performance0models covered
Safety9models covered
Human preference13models covered