Why this benchmark is useful
When you want a practical, industry-flavoured composite — finance, coding, and related work — as one Vals evidence row, not a VerdictPal endorsement.
Vals AI Index
A Vals composite score across industry-flavoured tasks such as finance and coding, intended to approximate practical model usefulness on economically relevant work.
When you want a practical, industry-flavoured composite — finance, coding, and related work — as one Vals evidence row, not a VerdictPal endorsement.
Percentage-style composite. Parts of the suite are proprietary — you cannot fully re-run it. A high index can hide uneven domain strengths; check the source page for current task mix.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Kimi K3Moonshot | 2026-07-16 | 74.7% | Vals AI — Vals Index2026-07-16 | Source-checked |
| 2 | Claude Mythos PreviewAnthropicResearch-preview row superseded by Claude Fable 5 (GA, 75.15% Vals Index, 2026-06-09). Mythos 5 ships to Glasswing partners only. | 2026-06-01 | 73.42% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 3 | GPT-5.6 SolOpenAI | 2026-07-09 | 73.1% | Vals AI — Vals Index2026-07-16 | Source-checked |
| 4 | Claude Opus 4.8Anthropic | 2026-05-28 | 70.4% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 5 | Claude Fable 5.1AnthropicVals Index key takeaway on 5 Sep 2026: Claude Fable 5.1 leads at 68.83%. The 1 Sep snapshot was 67.866%. | 2026-09-01 | 68.83% | Vals AI — Vals Index2026-09-05 | Source-checked |
| 6 | GPT 5.5OpenAI | 2026-04-23 | 68% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 7 | Claude Opus 5AnthropicVals Index 67.213% on the 1 Sep 2026 snapshot; rounded 67.21. Still #2 on the 5 Sep 2026 board behind Fable 5.1. | 2026-07-24 | 67.21% | Vals AI — Vals Index2026-09-01 | Source-checked |
| 8 | GPT-6 AstraOpenAIVals Index 66.61% on the 5 Sep 2026 board — #3 behind Fable 5.1 (68.83) and Opus 5 (67.21). Vals says Astra leads Terminal-Bench and Code Migration on that index. | 2026-09-03 | 66.61% | Vals AI — Vals Index2026-09-05 | Source-checked |
| 9 | Claude Opus 4.7Anthropic | 2026-04-16 | 66.1% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 10 | Claude Fable 5AnthropicVals Index 66.036% on the 1 Sep 2026 snapshot (rounded 66.04). Behind Fable 5.1 (67.87) and Opus 5 (67.21). Earlier 75.1 figure was a July 2026 board. | 2026-06-09 | 66.04% | Vals AI — Vals Index2026-09-01 | Source-checked |
| 11 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 60.3% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 12 | DeepSeek V4 ProDeepSeek | 2026-04-24 | 56.23% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 13 | Kimi K2.6 ThinkingMoonshot | 2026-04-20 | 55.55% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 14 | Gemini 3.1 Pro PreviewGoogle | 2026-02-19 | 53.42% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 15 | GPT-5.4 MiniOpenAI | 2026-03-17 | 51.42% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 16 | Gemini 3 Flash PreviewGoogle | 2025-12-17 | 49.31% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 17 | Qwen 3.7 MaxAlibaba | 2026-05-20 | 48.04% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 18 | Grok 4.3xAI | 2026-04-30 | 46.63% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 19 | GPT-5.4 NanoOpenAI | 2026-03-17 | 46.46% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 20 | MiniMax M2.7MiniMax | 2026-04-15 | 41.41% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 21 | Claude Haiku 4.5 ThinkingAnthropic | 2025-10-01 | 40.33% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 22 | Grok 4.20 0309 ReasoningxAI | 2026-03-05 | 39.11% | Vals AI — Vals Index2026-06-04 | Source-checked |
| 23 | Gemini 3.1 Flash Lite PreviewGoogle | 2026-05-07 | 35.24% | Vals AI — Vals Index2026-06-04 | Source-checked |
Top rows were refreshed from the public Vals index on 2026-07-19. Academic overlaps (SWE-bench, GPQA, etc.) stay on their official slugs; Vals runs are cross-linked in those explainers only.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
A Vals composite score across industry-flavoured tasks such as finance and coding, intended to approximate practical model usefulness on economically relevant work.
A strong Vals Index result says nothing about:
Percentage-style composite score. Read the source page for the current included tasks and weighting. Task format: Composite index built from Vals' industry benchmark suite; public pages expose model score, rank, and update date.
Vals Index is currently marked Active in the atlas.
Contamination risk for Vals Index is graded Unknown contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.