65%
AA-Omniscience Accuracy
Models / Anthropic
Claude Fable · Released 2026-06-09
Anthropic's June 2026 Mythos-class GA release. 1M context, 128K output, safety classifiers route high-risk queries to Opus 4.8. $10 / $50 per 1M tokens. Vals Index 66.04% on the 1 Sep 2026 snapshot — behind Fable 5.1 (67.87%) and Opus 5 (67.21%).
Long-horizon agentic work — multi-day coding migrations, finance analysis, vision-heavy rebuilds, and research pipelines that outlast Sonnet or Opus sessions.
Family profile
Best published score in each covered benchmark family.
5 tested benchmark families
Where Claude Fable 5 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.
The three highest-scoring models with pages in each capability family. Where Claude Fable 5 shows up, it's highlighted.
Benchmark rows added to the public ledger for Claude Fable 5 in the last 120 days. Older rows live in the full table below.
| Benchmark | Family | Score | Source | Days ago |
|---|---|---|---|---|
| Vals Index | Agentic | 66.04% | Vals AI — Vals Index | 5 |
| CritPt | Reasoning | 28.6% | Artificial Analysis — CritPt | 5 |
| AA-Omniscience | Knowledge | 65% | Artificial Analysis — AA-Omniscience | 5 |
| AA-LCR | Long context | 76.67% | Artificial Analysis | 5 |
| AA output speed | Performance | 58 t/s | Artificial Analysis | 5 |
| AA time to first token | Performance | 54.75s | Artificial Analysis | 5 |
| Artificial Analysis Intelligence Index | Reasoning | 62.1% | Artificial Analysis | 5 |
| GPQA Diamond | Reasoning | 92.6% | Artificial Analysis | 5 |
| Humanity's Last Exam | Reasoning | 55.5% | Artificial Analysis | 5 |
| IFBench | Reasoning | 63.47% | Artificial Analysis | 5 |
| SciCode | Coding | 60.2% | Artificial Analysis | 5 |
| τ³-Bench Banking | Agentic | 38.14% | Artificial Analysis | 5 |
| Terminal-Bench | Agentic | 84.64% | Artificial Analysis | 5 |
| Terminal-Bench-Science | Agentic | 21.4% | Terminal-Bench-Science 0.1 announcement | 10 |
| DeepSWE | Coding | 70% | DeepSWE leaderboard | 33 |
| GDPval | Agentic | 1818 | Artificial Analysis — GDPval-AA v2 | 81 |
A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.
Hover any dot for name, score, and input price. Keyboard: tab through the top twelve, or use the ranking below.
Every catalog benchmark for Claude Fable 5. Scores link to the original source; gaps mean no public row exists yet.
16 of 37 catalog benchmarks have a sourced row for Claude Fable 5.
65%
AA-Omniscience Accuracy
62.1%
AA Intelligence Index v4.1
28.6%
CritPt composite (AA run)
92.6%
GPQA Diamond (AA run)
55.5%
HLE (AA run)
63.47%
IFBench (AA run)
1818
GDPval-AA v2
84.64%
Terminal-Bench v2.1 (AA run)
21.4%
TB-Science 0.1 resolution (Claude Code)
66.04%
Vals Index
38.14%
τ³-Bench Banking (AA run)
76.67%
AA-LCR (AA run)
58 t/s
Median output tokens/s (1k prompt, default provider)
54.75s
Median time to first token (1k prompt, default provider)