Benchmarks / Agentic

Vals Index

Vals AI Index

A Vals composite score across industry-flavoured tasks such as finance and coding, intended to approximate practical model usefulness on economically relevant work.

What this does not measure
  • Open reproducibility — parts of the benchmark use proprietary, non-public datasets, so readers cannot fully re-run the suite.
  • A single capability — the index blends domains, so a high score can hide uneven strengths across finance, coding, education, or multimodal tasks.
  • VerdictPal endorsement — Vals is useful evidence, but it remains one source row with its own methodology and weighting choices.
Analysis

Why this benchmark is useful

When you want a practical, industry-flavoured composite — finance, coding, and related work — as one Vals evidence row, not a VerdictPal endorsement.

Scope

Coverage map

Task family
Agentic
Format
Composite index built from Vals' industry benchmark suite; public pages expose model score, rank, and update date.
Scoring
Percentage-style composite score. Read the source page for the current included tasks and weighting.
Maintainer
Vals AI
Reading guide

How to read the scores

Percentage-style composite. Parts of the suite are proprietary — you cannot fully re-run it. A high index can hide uneven domain strengths; check the source page for current task mix.

Blind spots

What it does not cover

  • Open reproducibility — parts of the benchmark use proprietary, non-public datasets, so readers cannot fully re-run the suite.
  • A single capability — the index blends domains, so a high score can hide uneven strengths across finance, coding, education, or multimodal tasks.
  • VerdictPal endorsement — Vals is useful evidence, but it remains one source row with its own methodology and weighting choices.
Scores

Evidence ledger

23 rows
2323 rows
74.7%best score
23source-checked
1sources
2026-09-05to 2026-06-04
Vals AI23

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Fable 5.168.83% · Vals AI — Vals Index
  2. GPT-6 Astra66.61% · Vals AI — Vals Index
  3. Claude Opus 567.21% · Vals AI — Vals Index
  4. Claude Fable 566.04% · Vals AI — Vals Index
  5. Kimi K374.7% · Vals AI — Vals Index

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Kimi K3Moonshot2026-07-1674.7%
Vals AI — Vals Index2026-07-16
Source-checked
2Claude Mythos PreviewAnthropicResearch-preview row superseded by Claude Fable 5 (GA, 75.15% Vals Index, 2026-06-09). Mythos 5 ships to Glasswing partners only.2026-06-0173.42%
Vals AI — Vals Index2026-06-04
Source-checked
3GPT-5.6 SolOpenAI2026-07-0973.1%
Vals AI — Vals Index2026-07-16
Source-checked
4Claude Opus 4.8Anthropic2026-05-2870.4%
Vals AI — Vals Index2026-06-04
Source-checked
5Claude Fable 5.1AnthropicVals Index key takeaway on 5 Sep 2026: Claude Fable 5.1 leads at 68.83%. The 1 Sep snapshot was 67.866%.2026-09-0168.83%
Vals AI — Vals Index2026-09-05
Source-checked
6GPT 5.5OpenAI2026-04-2368%
Vals AI — Vals Index2026-06-04
Source-checked
7Claude Opus 5AnthropicVals Index 67.213% on the 1 Sep 2026 snapshot; rounded 67.21. Still #2 on the 5 Sep 2026 board behind Fable 5.1.2026-07-2467.21%
Vals AI — Vals Index2026-09-01
Source-checked
8GPT-6 AstraOpenAIVals Index 66.61% on the 5 Sep 2026 board — #3 behind Fable 5.1 (68.83) and Opus 5 (67.21). Vals says Astra leads Terminal-Bench and Code Migration on that index.2026-09-0366.61%
Vals AI — Vals Index2026-09-05
Source-checked
9Claude Opus 4.7Anthropic2026-04-1666.1%
Vals AI — Vals Index2026-06-04
Source-checked
10Claude Fable 5AnthropicVals Index 66.036% on the 1 Sep 2026 snapshot (rounded 66.04). Behind Fable 5.1 (67.87) and Opus 5 (67.21). Earlier 75.1 figure was a July 2026 board.2026-06-0966.04%
Vals AI — Vals Index2026-09-01
Source-checked
11Claude Sonnet 4.6Anthropic2026-02-1760.3%
Vals AI — Vals Index2026-06-04
Source-checked
12DeepSeek V4 ProDeepSeek2026-04-2456.23%
Vals AI — Vals Index2026-06-04
Source-checked
13Kimi K2.6 ThinkingMoonshot2026-04-2055.55%
Vals AI — Vals Index2026-06-04
Source-checked
14Gemini 3.1 Pro PreviewGoogle2026-02-1953.42%
Vals AI — Vals Index2026-06-04
Source-checked
15GPT-5.4 MiniOpenAI2026-03-1751.42%
Vals AI — Vals Index2026-06-04
Source-checked
16Gemini 3 Flash PreviewGoogle2025-12-1749.31%
Vals AI — Vals Index2026-06-04
Source-checked
17Qwen 3.7 MaxAlibaba2026-05-2048.04%
Vals AI — Vals Index2026-06-04
Source-checked
18Grok 4.3xAI2026-04-3046.63%
Vals AI — Vals Index2026-06-04
Source-checked
19GPT-5.4 NanoOpenAI2026-03-1746.46%
Vals AI — Vals Index2026-06-04
Source-checked
20MiniMax M2.7MiniMax2026-04-1541.41%
Vals AI — Vals Index2026-06-04
Source-checked
21Claude Haiku 4.5 ThinkingAnthropic2025-10-0140.33%
Vals AI — Vals Index2026-06-04
Source-checked
22Grok 4.20 0309 ReasoningxAI2026-03-0539.11%
Vals AI — Vals Index2026-06-04
Source-checked
23Gemini 3.1 Flash Lite PreviewGoogle2026-05-0735.24%
Vals AI — Vals Index2026-06-04
Source-checked
Method

What it covers

Top rows were refreshed from the public Vals index on 2026-07-19. Academic overlaps (SWE-bench, GPQA, etc.) stay on their official slugs; Vals runs are cross-linked in those explainers only.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Tools that report it

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does Vals Index measure?

A Vals composite score across industry-flavoured tasks such as finance and coding, intended to approximate practical model usefulness on economically relevant work.

What does a high Vals Index score not prove?

A strong Vals Index result says nothing about:

  • Open reproducibility — parts of the benchmark use proprietary, non-public datasets, so readers cannot fully re-run the suite.
  • A single capability — the index blends domains, so a high score can hide uneven strengths across finance, coding, education, or multimodal tasks.
  • VerdictPal endorsement — Vals is useful evidence, but it remains one source row with its own methodology and weighting choices.

How is Vals Index scored?

Percentage-style composite score. Read the source page for the current included tasks and weighting. Task format: Composite index built from Vals' industry benchmark suite; public pages expose model score, rank, and update date.

Is Vals Index saturated?

Vals Index is currently marked Active in the atlas.

Can Vals Index results be contaminated by training data?

Contamination risk for Vals Index is graded Unknown contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.