Why this benchmark is useful
It tests whether a system can act through tools over several steps. Use it when you care about execution, recovery, and task completion instead of chat fluency.
Benchmarks / Agentic
Vals AI Index
A Vals composite score across industry-flavoured tasks such as finance and coding, intended to approximate practical model usefulness on economically relevant work.
It tests whether a system can act through tools over several steps. Use it when you care about execution, recovery, and task completion instead of chat fluency.
Read Vals Index as a active signal with unknown contamination risk. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Fable 5AnthropicVals reports 75.14% ± 0.64 on the public model page; index copy rounds to 75.1. Snapshot refreshed July 19, 2026. | 2026-06-09 | 75.1% | Vals AI — Vals Index2026-07-16 | Source-checked |
| 2 | Kimi K3Moonshot | 2026-07-16 | 74.7% | Vals AI — Vals Index2026-07-16 | Source-checked |
| 3 | Claude Mythos PreviewAnthropicResearch-preview row superseded by Claude Fable 5 (GA, 75.15% Vals Index, 2026-06-09). Mythos 5 ships to Glasswing partners only. | 2026-06-01 | 73.42% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 4 | GPT-5.6 SolOpenAI | 2026-07-09 | 73.1% | Vals AI — Vals Index2026-07-16 | Source-checked |
| 5 | Claude Opus 4.8Anthropic | 2026-05-28 | 70.4% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 6 | GPT 5.5OpenAI | 2026-04-23 | 68% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 7 | Claude Opus 4.7Anthropic | 2026-04-16 | 66.1% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 8 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 60.3% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 9 | DeepSeek V4 ProDeepSeek | 2026-04-24 | 56.23% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 10 | Kimi K2.6 ThinkingMoonshot | 2026-04-20 | 55.55% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 11 | Gemini 3.1 Pro PreviewGoogle | 2026-02-19 | 53.42% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 12 | GPT-5.4 MiniOpenAI | 2026-03-17 | 51.42% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 13 | Gemini 3 Flash PreviewGoogle | 2025-12-17 | 49.31% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 14 | Qwen 3.7 MaxAlibaba | 2026-05-20 | 48.04% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 15 | Grok 4.3xAI | 2026-04-30 | 46.63% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 16 | GPT-5.4 NanoOpenAI | 2026-03-17 | 46.46% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 17 | MiniMax M2.7MiniMax | 2026-04-15 | 41.41% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 18 | Claude Haiku 4.5 ThinkingAnthropic | 2025-10-01 | 40.33% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 19 | Grok 4.20 0309 ReasoningxAI | 2026-03-05 | 39.11% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 20 | Gemini 3.1 Flash Lite PreviewGoogle | 2026-05-07 | 35.24% | Vals AI — Vals Index2026-06-04 | Source-checked |
Top rows were refreshed from the public Vals index on 2026-07-19. Academic overlaps (SWE-bench, GPQA, etc.) stay on their official slugs; Vals runs are cross-linked in those explainers only.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
The Pack · Editorial newsletter
One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.