Benchmarks / Agentic

Vals Index

Vals AI Index

A Vals composite score across industry-flavoured tasks such as finance and coding, intended to approximate practical model usefulness on economically relevant work.

AgenticSolidActiveUnknown contamination riskSince 2026
What this does not measure
  • Open reproducibility — parts of the benchmark use proprietary, non-public datasets, so readers cannot fully re-run the suite.
  • A single capability — the index blends domains, so a high score can hide uneven strengths across finance, coding, education, or multimodal tasks.
  • VerdictPal endorsement — Vals is useful evidence, but it remains one source row with its own methodology and weighting choices.
Analysis

Why this benchmark is useful

It tests whether a system can act through tools over several steps. Use it when you care about execution, recovery, and task completion instead of chat fluency.

Scope

Coverage map

Task family
Agentic
Format
Composite index built from Vals' industry benchmark suite; public pages expose model score, rank, and update date.
Scoring
Percentage-style composite score. Read the source page for the current included tasks and weighting.
Maintainer
Vals AI
Reading guide

How to read the scores

Read Vals Index as a active signal with unknown contamination risk. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.

Blind spots

What it does not cover

  • Open reproducibility — parts of the benchmark use proprietary, non-public datasets, so readers cannot fully re-run the suite.
  • A single capability — the index blends domains, so a high score can hide uneven strengths across finance, coding, education, or multimodal tasks.
  • VerdictPal endorsement — Vals is useful evidence, but it remains one source row with its own methodology and weighting choices.
Scores

Evidence ledger

20 rows
2020 rows
75.1%best score
20source-checked
1sources
2026-07-16to 2026-06-04
Vals AI20

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Fable 575.1% · Vals AI — Vals Index
  2. Kimi K374.7% · Vals AI — Vals Index
  3. GPT-5.6 Sol73.1% · Vals AI — Vals Index
  4. Claude Mythos Preview73.42% · Vals AI — Vals Index
  5. Claude Opus 4.870.4% · Vals AI — Vals Index

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Fable 5AnthropicVals reports 75.14% ± 0.64 on the public model page; index copy rounds to 75.1. Snapshot refreshed July 19, 2026.2026-06-0975.1%
Vals AI — Vals Index2026-07-16
Source-checked
2Kimi K3Moonshot2026-07-1674.7%
Vals AI — Vals Index2026-07-16
Source-checked
3Claude Mythos PreviewAnthropicResearch-preview row superseded by Claude Fable 5 (GA, 75.15% Vals Index, 2026-06-09). Mythos 5 ships to Glasswing partners only.2026-06-0173.42%
Vals AI — Vals Index2026-06-04
Source-checked
4GPT-5.6 SolOpenAI2026-07-0973.1%
Vals AI — Vals Index2026-07-16
Source-checked
5Claude Opus 4.8Anthropic2026-05-2870.4%
Vals AI — Vals Index2026-06-04
Source-checked
6GPT 5.5OpenAI2026-04-2368%
Vals AI — Vals Index2026-06-04
Source-checked
7Claude Opus 4.7Anthropic2026-04-1666.1%
Vals AI — Vals Index2026-06-04
Source-checked
8Claude Sonnet 4.6Anthropic2026-02-1760.3%
Vals AI — Vals Index2026-06-04
Source-checked
9DeepSeek V4 ProDeepSeek2026-04-2456.23%
Vals AI — Vals Index2026-06-04
Source-checked
10Kimi K2.6 ThinkingMoonshot2026-04-2055.55%
Vals AI — Vals Index2026-06-04
Source-checked
11Gemini 3.1 Pro PreviewGoogle2026-02-1953.42%
Vals AI — Vals Index2026-06-04
Source-checked
12GPT-5.4 MiniOpenAI2026-03-1751.42%
Vals AI — Vals Index2026-06-04
Source-checked
13Gemini 3 Flash PreviewGoogle2025-12-1749.31%
Vals AI — Vals Index2026-06-04
Source-checked
14Qwen 3.7 MaxAlibaba2026-05-2048.04%
Vals AI — Vals Index2026-06-04
Source-checked
15Grok 4.3xAI2026-04-3046.63%
Vals AI — Vals Index2026-06-04
Source-checked
16GPT-5.4 NanoOpenAI2026-03-1746.46%
Vals AI — Vals Index2026-06-04
Source-checked
17MiniMax M2.7MiniMax2026-04-1541.41%
Vals AI — Vals Index2026-06-04
Source-checked
18Claude Haiku 4.5 ThinkingAnthropic2025-10-0140.33%
Vals AI — Vals Index2026-06-04
Source-checked
19Grok 4.20 0309 ReasoningxAI2026-03-0539.11%
Vals AI — Vals Index2026-06-04
Source-checked
20Gemini 3.1 Flash Lite PreviewGoogle2026-05-0735.24%
Vals AI — Vals Index2026-06-04
Source-checked
Method

What it covers

Top rows were refreshed from the public Vals index on 2026-07-19. Academic overlaps (SWE-bench, GPQA, etc.) stay on their official slugs; Vals runs are cross-linked in those explainers only.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Tools that report it

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.