Why this benchmark is useful
For source-heavy students the failure mode that matters is confident wrong answers. AA-Omniscience is the public suite that jointly scores recall and hallucination across domains that actually show up in papers.
AA-Omniscience
Factual recall and knowledge calibration across 6,000 questions in 42 economically relevant topics and six domains. Accuracy is 8% of AA Intelligence Index v4.1; the non-hallucination slice is another 4%.
For source-heavy students the failure mode that matters is confident wrong answers. AA-Omniscience is the public suite that jointly scores recall and hallucination across domains that actually show up in papers.
Read Accuracy as the headline, then the hallucination rate in the notes. Claude Fable 5.1 leads the 1 Sep 2026 AA board on both Accuracy (~67%) and Index (43). GPT-5.6 Sol is close on Accuracy (~59%) with a much higher hallucination rate. Do not mix Index and Accuracy in one ranking.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Fable 5.1AnthropicAA public Omniscience board, 1 Sep 2026. Headline Accuracy 67%; companion Index 43. Do not rank Index against Accuracy. | 2026-09-01 | 67% | 2026-09-01 | Source-checked |
| 2 | Claude Fable 5AnthropicAA headline Accuracy 65% (Opus 4.8 fallback variant); Index 43. | 2026-06-09 | 65% | 2026-09-01 | Source-checked |
| 3 | Claude Opus 5Anthropic | 2026-07-24 | 60.9% | 2026-09-01 | Source-checked |
| 4 | GPT-5.6 SolOpenAIHigh Accuracy with elevated hallucination vs Claude on this suite. Index is lower than Accuracy implies. | 2026-07-09 | 59.4% | 2026-09-01 | Source-checked |
| 5 | GPT-5.5OpenAI | 2026-04-23 | 57% | 2026-09-01 | Source-checked |
| 11 | Gemini 3.5 FlashGoogle | 2026-05-19 | 51.9% | 2026-09-01 | Source-checked |
| 12 | Grok 4.5xAI | 2026-07-08 | 51.6% | 2026-09-01 | Source-checked |
| 16 | DeepSeek V4 ProDeepSeek | 2026-08-13 | 49.1% | 2026-09-01 | Source-checked |
| 18 | Claude Opus 4.8Anthropic | 2026-05-28 | 48.8% | 2026-09-01 | Source-checked |
| 19 | Grok 4.6xAI | 2026-08-12 | 48.2% | 2026-09-01 | Source-checked |
| 20 | Kimi K3Moonshot | 2026-07-16 | 47.6% | 2026-09-01 | Source-checked |
| 22 | GPT-5.6 TerraOpenAI | 2026-07-09 | 46.8% | 2026-09-01 | Source-checked |
Accuracy rows ingested from Artificial Analysis' public Omniscience evaluation (1 Sep 2026 board). A public Hugging Face subset exists; the live AA board is the source of these scores. The original 2025 paper had almost no models above Index 0 — 2026 Claude rows now sit in the 30–43 Index band.
Accuracy is not saturated — 2026 frontier rows cluster in the 45–67% band. The Index remains well below 50.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
Factual recall and knowledge calibration across 6,000 questions in 42 economically relevant topics and six domains. Accuracy is 8% of AA Intelligence Index v4.1; the non-hallucination slice is another 4%.
A strong AA-Omniscience result says nothing about:
VerdictPal ledger rows are AA-Omniscience Accuracy (% of all questions answered correctly, including abstentions as misses). The companion Index (−100 to +100) rewards correct answers, penalizes hallucinations, and scores 0 for abstention — we keep Index figures in notes, not as the sortable score, because the Index can go negative. Task format: 6,000 questions from academic and industry sources across business, humanities, science/engineering, health, law, and software engineering. Models may answer or abstain.
AA-Omniscience is currently marked Active in the atlas. Ceiling context: Accuracy is not saturated — 2026 frontier rows cluster in the 45–67% band. The Index remains well below 50.
Contamination risk for AA-Omniscience is graded Medium contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.