Benchmarks / Knowledge

Omniscience Accuracy

AA-Omniscience

Factual recall and knowledge calibration across 6,000 questions in 42 economically relevant topics and six domains. Accuracy is 8% of AA Intelligence Index v4.1; the non-hallucination slice is another 4%.

What this does not measure
  • Open-web research skill — questions are closed-book. A model that should search and cite is not being tested here.
  • Whether a refusal is useful — abstention scores 0 on the Index; it is better than a confident wrong answer, not a research deliverable.
  • A single reliability number without the hallucination rate — high accuracy with high hallucination is the usual OpenAI-shaped failure on this suite.
Analysis

Why this benchmark is useful

For source-heavy students the failure mode that matters is confident wrong answers. AA-Omniscience is the public suite that jointly scores recall and hallucination across domains that actually show up in papers.

Scope

Coverage map

Task family
Knowledge
Format
6,000 questions from academic and industry sources across business, humanities, science/engineering, health, law, and software engineering. Models may answer or abstain.
Scoring
VerdictPal ledger rows are AA-Omniscience Accuracy (% of all questions answered correctly, including abstentions as misses). The companion Index (−100 to +100) rewards correct answers, penalizes hallucinations, and scores 0 for abstention — we keep Index figures in notes, not as the sortable score, because the Index can go negative.
Maintainer
Artificial Analysis
Reading guide

How to read the scores

Read Accuracy as the headline, then the hallucination rate in the notes. Claude Fable 5.1 leads the 1 Sep 2026 AA board on both Accuracy (~67%) and Index (43). GPT-5.6 Sol is close on Accuracy (~59%) with a much higher hallucination rate. Do not mix Index and Accuracy in one ranking.

Blind spots

What it does not cover

  • Open-web research skill — questions are closed-book. A model that should search and cite is not being tested here.
  • Whether a refusal is useful — abstention scores 0 on the Index; it is better than a confident wrong answer, not a research deliverable.
  • A single reliability number without the hallucination rate — high accuracy with high hallucination is the usual OpenAI-shaped failure on this suite.
Scores

Evidence ledger

12 rows
1212 rows
67%best score
12source-checked
1sources
2026-09-01source date
12

api · api

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Fable 5.167% ·
  2. Claude Fable 565% ·
  3. Claude Opus 560.9% ·
  4. GPT-5.6 Sol59.4% ·
  5. GPT-5.557% ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Fable 5.1AnthropicAA public Omniscience board, 1 Sep 2026. Headline Accuracy 67%; companion Index 43. Do not rank Index against Accuracy.2026-09-0167%
2026-09-01
Source-checked
2Claude Fable 5AnthropicAA headline Accuracy 65% (Opus 4.8 fallback variant); Index 43.2026-06-0965%
2026-09-01
Source-checked
3Claude Opus 5Anthropic2026-07-2460.9%
2026-09-01
Source-checked
4GPT-5.6 SolOpenAIHigh Accuracy with elevated hallucination vs Claude on this suite. Index is lower than Accuracy implies.2026-07-0959.4%
2026-09-01
Source-checked
5GPT-5.5OpenAI2026-04-2357%
2026-09-01
Source-checked
11Gemini 3.5 FlashGoogle2026-05-1951.9%
2026-09-01
Source-checked
12Grok 4.5xAI2026-07-0851.6%
2026-09-01
Source-checked
16DeepSeek V4 ProDeepSeek2026-08-1349.1%
2026-09-01
Source-checked
18Claude Opus 4.8Anthropic2026-05-2848.8%
2026-09-01
Source-checked
19Grok 4.6xAI2026-08-1248.2%
2026-09-01
Source-checked
20Kimi K3Moonshot2026-07-1647.6%
2026-09-01
Source-checked
22GPT-5.6 TerraOpenAI2026-07-0946.8%
2026-09-01
Source-checked
Method

What it covers

Accuracy rows ingested from Artificial Analysis' public Omniscience evaluation (1 Sep 2026 board). A public Hugging Face subset exists; the live AA board is the source of these scores. The original 2025 paper had almost no models above Index 0 — 2026 Claude rows now sit in the 30–43 Index band.

Score ceiling

Where it breaks down

Accuracy is not saturated — 2026 frontier rows cluster in the 45–67% band. The Index remains well below 50.

Tools that report it

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does AA-Omniscience measure?

Factual recall and knowledge calibration across 6,000 questions in 42 economically relevant topics and six domains. Accuracy is 8% of AA Intelligence Index v4.1; the non-hallucination slice is another 4%.

What does a high AA-Omniscience score not prove?

A strong AA-Omniscience result says nothing about:

  • Open-web research skill — questions are closed-book. A model that should search and cite is not being tested here.
  • Whether a refusal is useful — abstention scores 0 on the Index; it is better than a confident wrong answer, not a research deliverable.
  • A single reliability number without the hallucination rate — high accuracy with high hallucination is the usual OpenAI-shaped failure on this suite.

How is AA-Omniscience scored?

VerdictPal ledger rows are AA-Omniscience Accuracy (% of all questions answered correctly, including abstentions as misses). The companion Index (−100 to +100) rewards correct answers, penalizes hallucinations, and scores 0 for abstention — we keep Index figures in notes, not as the sortable score, because the Index can go negative. Task format: 6,000 questions from academic and industry sources across business, humanities, science/engineering, health, law, and software engineering. Models may answer or abstain.

Is AA-Omniscience saturated?

AA-Omniscience is currently marked Active in the atlas. Ceiling context: Accuracy is not saturated — 2026 frontier rows cluster in the 45–67% band. The Index remains well below 50.

Can AA-Omniscience results be contaminated by training data?

Contamination risk for AA-Omniscience is graded Medium contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.