VerdictPal · editorial desk · 2026VerdictPal
Models

Model-first evidence atlas.

We score models, not vendors. Below: every carded model the public benchmark ledger has evidence for, sorted by intelligence receipts (Vals Index first, then Artificial Analysis). Every row carries a date and a URL, not a universal verdict.
Current Vals evidence anchor · Vals Index
Claude Fable 5Anthropic · Claude Fable · 75 on the Vals intelligence index
111
Carded models
21
Providers
85
Priced in USD/1M
5
In score ledger only

Vals Index rows appear first. Models without a Vals row fall back to Artificial Analysis. Source badges show whether a row is single-source or corroborated — not a cross-benchmark winner.

  1. 1
    Claude Fable 5Anthropic · Claude Fable · Vals + AA agree
    75Vals
  2. 2
    Kimi K3Moonshot · Kimi K3 · Vals + AA agree
    75Vals
  3. 3
    Claude Mythos PreviewAnthropic · Claude Opus · Vals only
    73Vals
  4. 4
    GPT-5.6 SolOpenAI · GPT-5.6 · Vals + AA agree
    73Vals
  5. 5
    Claude Opus 4.8Anthropic · Claude Opus · Vals + AA agree
    70Vals
  6. 6
    GPT 5.5OpenAI · GPT 5 · Vals + AA agree
    68Vals
  7. 7
    Claude Opus 4.7Anthropic · Claude Opus · Vals + AA agree
    66Vals
  8. 8
    Claude Sonnet 4.6Anthropic · Claude Sonnet · Vals + AA agree
    60Vals
  9. 9
    DeepSeek V4 Pro MaxDeepSeek · DeepSeek · Vals only
    56Vals
  10. 10
    Kimi K2.6Moonshot · Kimi · Vals + AA agree
    56Vals
  11. 11
    Gemini 3.1 ProGoogle · Gemini · Vals only
    53Vals
  12. 12
    GPT-5.4 MiniOpenAI · GPT-5.4 · Vals + AA agree
    51Vals
  13. 13
    Gemini 3.5 FlashGoogle · Gemini · Vals + AA agree
    49Vals
  14. 14
    Gemini 3 Flash PreviewGoogle · Gemini 3 · Vals + AA agree
    49Vals
  15. 15
    Qwen 3.7 MaxAlibaba · Qwen · Vals only
    48Vals
  16. 16
    Grok 4.3xAI · Grok · Vals + AA agree
    47Vals
  17. 17
    GPT-5.4 NanoOpenAI · GPT-5.4 · Vals + AA agree
    46Vals
  18. 18
    MiniMax M2.7MiniMax · MiniMax · Vals + AA agree
    41Vals
  19. 19
    Claude Haiku 4.5Anthropic · Claude Haiku · Vals only
    40Vals
  20. 20
    Grok 4.20 ReasoningxAI · Grok 4 · Vals + AA agree
    39Vals
  21. 21
    Gemini 3.1 Flash LiteGoogle · Gemini 3.1 · Vals + AA agree
    35Vals
  22. 22
    GPT-5.6 TerraOpenAI · GPT-5.6 · AA only
    55AA
  23. 23
    Grok 4.5xAI · Grok · AA only
    54AA
  24. 24
    GPT-5.5 (high)OpenAI · GPT-5.5 (high) · AA only
    53AA
  25. 25
    Claude Sonnet 5Anthropic · Claude Sonnet · AA only
    53AA
  26. 26
    GPT 5.4OpenAI · GPT 5 · AA only
    51AA
  27. 27
    GPT-5.6 LunaOpenAI · GPT-5.6 · AA only
    51AA
  28. 28
    Muse Spark 1.1Meta · Muse Spark · AA only
    51AA
  29. 29
    GLM-5.2Zhipu · GLM · AA only
    51AA
  30. 30
    GPT-5.5 (medium)OpenAI · GPT-5.5 (medium) · AA only
    50AA
  31. 31
    Gemini 3.1 Pro PreviewGoogle · Gemini 3.1 · AA only
    47AA
  32. 32
    Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic · Claude Sonnet · AA only
    47AA
  33. 33
    Qwen3.7 MaxAlibaba · Qwen3.7 Max · AA only
    46AA
  34. 34
    Gemini 3.5 Flash (medium)Google · Gemini 3.5 · AA only
    45AA
  35. 35
    GPT-5.3 Codex (xhigh)OpenAI · GPT-5.3 Codex · AA only
    44AA
  36. 36
    MiniMax M3MiniMax · MiniMax · AA only
    44AA
  37. 37
    Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic · Claude Opus · AA only
    44AA
  38. 38
    DeepSeek V4 Pro (Reasoning, Max Effort)DeepSeek · DeepSeek V4 · AA only
    44AA
  39. 39
    GPT-5.5 (low)OpenAI · GPT-5.5 (low) · AA only
    44AA
  40. 40
    Muse SparkOther · Muse · AA only
    43AA
  41. 41
    Claude Opus 4.7 (Non-reasoning, High Effort)Anthropic · Claude Opus · AA only
    43AA
  42. 42
    GPT-5.2OpenAI · GPT-5 · AA only
    42AA
  43. 43
    MiMo-V2.5-ProXiaomi · Xiaomi · AA only
    42AA
  44. 44
    Kimi K2.7 CodeMoonshot · Kimi · AA only
    42AA
  45. 45
    Gemini 3 ProGoogle · Gemini 3 · AA only
    40AA
  46. 46
    GLM-5.1Zhipu · GLM · AA only
    40AA
  47. 47
    Qwen 3.7 PlusAlibaba · Qwen · AA only
    39AA
  48. 48
    Claude Opus 4.6Anthropic · Claude Opus · AA only
    38AA
  49. 49
    Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA · Nemotron · AA only
    38AA
  50. 50
    GPT-5.1OpenAI · GPT-5 · AA only
    37AA
  51. 51
    Claude Opus 4.5Anthropic · Claude Opus · AA only
    35AA
  52. 52
    Kimi K2.5Moonshot · Kimi · AA only
    35AA
  53. 53
    MiniMax M2.5MiniMax · MiniMax · AA only
    34AA
  54. 54
    LongCat-2.0Meituan · LongCat · AA only
    34AA
  55. 55
    OpenAI o3OpenAI · OpenAI o · AA only
    30AA
  56. 56
    Step 3.7 FlashStepFun · Step · AA only
    30AA
  57. 57
    Gemma 4 31B ITGoogle · Gemma · AA only
    29AA
  58. 58
    OpenAI o1OpenAI · OpenAI o · AA only
    23AA
  59. 59
    Gemma 4 12B (Reasoning)Google · Gemma · AA only
    22AA
  60. 60
    DeepSeek R1DeepSeek · DeepSeek R1 · AA only
    20AA
  61. 61
    North Mini CodeCohere · North · AA only
    20AA
  62. 62
    HyperNova 60B 2605Multiverse · HyperNova · AA only
    18AA
  63. 63
    Mistral Large 3Mistral · Mistral Large · AA only
    16AA
  64. 64
    Llama 4 MaverickMeta · Llama 4 · AA only
    14AA
  65. 65
    Gemma 4 12B (Non-reasoning)Google · Gemma · AA only
    13AA
  66. 66
    Claude 3 OpusAnthropic · Claude Opus · AA only
    12AA
  67. 67
    MiniCPM5-1B (Reasoning)OpenBMB · MiniCPM · AA only
    12AA
  68. 68
    Llama 4 ScoutMeta · Llama 4 · AA only
    10AA
  69. 69
    Command ACohere · Command · AA only
    8AA
  70. 70
    LFM2.5-8B-A1BLiquid · LFM · AA only
    8AA
  71. 71
    DeepSeek Coder V2DeepSeek · DeepSeek Coder · AA only
    5AA
  72. 72
    Gemini 1.0 UltraGoogle · Gemini 1 · AA only
    5AA

Blended cost per million tokens (input + 3× output) on the X axis, log scale. Intelligence index on the Y axis — Vals when present, otherwise Artificial Analysis. Dots are sized by context window and connected by the Pareto frontier (best intelligence at each cost).

Pareto frontier · curated models53 scored · log scale
$0.10$1$10$1000255075100Blended $/1M (log)IntelligenceQwen 3.6 Plus
All 53 models in this chart
ModelIntelligenceBlended $/1MFrontier
GPT-5.4 Pro83$142.50Yes
GPT-5.5 Pro82$142.50
Qwen 3.6 Plus79$1.63Yes
DeepSeek V4 Flash Max78$0.25Yes
Kimi K375$12.00
Claude Fable 575$40.00
GPT-5.6 Sol73$23.75
Claude Mythos Preview73$243.75
Claude Opus 4.870$60.00
GPT 5.568$40.63
Claude Opus 4.766$20.00
Claude Sonnet 4.660$12.00
DeepSeek V4 Pro Max56$1.30
Kimi K2.656$2.02
GPT-5.6 Terra55$11.88
Grok 4.554$5.00
Claude Sonnet 553$8.00
Gemini 3.1 Pro53$17.50
GPT-5.5 (high)53$23.75
Muse Spark 1.151$3.50
GPT-5.4 Mini51$3.56
GLM-5.251$3.65
GPT-5.6 Luna51$4.75
GPT 5.451$22.75
GPT-5.5 (medium)50$23.75
Gemini 3.5 Flash49$0.97
Gemini 3 Flash Preview49$2.38
Qwen 3.7 Max48$1.00
Gemini 3.1 Pro Preview47$9.50
Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)47$12.00
Grok 4.347$20.00
GPT-5.4 Nano46$0.99
Qwen3.7 Max46$6.25
Gemini 3.5 Flash (medium)45$7.13
DeepSeek V4 Pro (Reasoning, Max Effort)44$0.76
MiniMax M344$1.60
GPT-5.3 Codex (xhigh)44$10.94
Claude Opus 4.6 (Adaptive Reasoning, Max Effort)44$20.00
GPT-5.5 (low)44$23.75
Claude Opus 4.7 (Non-reasoning, High Effort)43$20.00
MiMo-V2.5-Pro42$0.76
Kimi K2.7 Code42$3.24
GPT-5.242$10.94
MiniMax M2.741$1.21
Claude Haiku 4.540$4.00
Claude Opus 4.638$20.00
Gemini 3.1 Flash Lite35$1.19
Kimi K2.535$2.40
Claude Opus 4.535$20.00
MiniMax M2.534$0.97
LongCat-2.034$2.40
Mistral Large 316$1.25
Command A8$8.13

The earlier cut, using input cost alone. Useful when your workload is input-heavy (RAG, document Q&A) rather than output-heavy.

0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMiniMaxxAIMetaZhipuXiaomiMeituanCohere

Older releases, special configurations (Codex CLI, ForgeCode), and ad-hoc aggregations. We surface their sourced scores but we don't promote them to the model atlas until they earn a full model card.

The highest intelligence receipt in each family slice. Use this as a sanity check that a single number isn't hiding a weak spot.

Claude Fable 575

Anthropic · Claude Fable

ReasoningCodingAgenticLong context
Reasoning
93
Coding
70
Agentic
100
Long context
70

83average across 4 tested families

Kimi K375

Moonshot · Kimi K3

ReasoningCodingAgenticLong context
Reasoning
94
Coding
59
Agentic
85
Long context
75

78average across 4 tested families

Claude Mythos Preview73

Anthropic · Claude Opus

ReasoningCodingAgentic
Reasoning
95
Coding
94
Agentic
84

91average across 3 tested families

A model with no row in a family is a model that hasn't been publicly tested there. We surface that instead of pretending it's weak.

Open the benchmark atlas
Knowledge30models covered
Reasoning74models covered
Math25models covered
Coding74models covered
Agentic71models covered
Multimodal8models covered
Long context69models covered
Tool use20models covered
Performance0models covered
Safety9models covered
Human preference13models covered