Benchmarks / Knowledge

Citation Fidelity (VerdictPal Lab)

When an answer-engine cites a source, does the source actually say what the answer claims? We grade real research questions on whether every cited claim is supported by the linked page — the failure mode students get burned by most.

KnowledgeDraftActiveLow contamination riskVerdictPal LabSince 2026
What this does not measure
  • Answer quality or fluency — a beautifully written answer with a fabricated citation fails here on purpose.
  • Coverage — we test whether what is cited is true, not whether the model found every relevant source.
  • It is a small, hand-graded set — treat early numbers as directional, not definitive.
Analysis

Why this benchmark is useful

It gives a compact signal for a specific capability. Use it as one dated receipt beside pricing, privacy, and hands-on evidence.

Scope

Coverage map

Task family
Knowledge
Format
A fixed bank of 12 real student/researcher questions (policy brief, product comparison, academic-adjacent, local/current), run through each answer-engine; every cited claim is checked against its linked source by two reviewers.
Scoring
Citation-supported rate = supported cited claims / total cited claims, plus counts for unsupported and broken-link claims.
Maintainer
VerdictPal desk
Reading guide

How to read the scores

Read Citation Fidelity (VerdictPal Lab) as a active signal with low contamination risk. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.

Blind spots

What it does not cover

  • Answer quality or fluency — a beautifully written answer with a fabricated citation fails here on purpose.
  • Coverage — we test whether what is cited is true, not whether the model found every relevant source.
  • It is a small, hand-graded set — treat early numbers as directional, not definitive.
Scores

Evidence ledger

0 rows

No results published yet.

The protocol is public before the run ships.

Method

What it covers

Protocol v0.1 frozen. July 2026 desk run: 156/156 capture slots filled across 12 questions and 13 engines — human grading and sign-off pending before any score ships on the atlas.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

How we run it

Frozen protocol v0.1 (2026-07-06). Twelve questions (Q01–Q12): three policy brief, three product comparison, three academic-adjacent, three local/current. Each engine gets the same prompt and default settings; two reviewers grade every cited claim independently (supported / unsupported / broken-link); disagreements are resolved on the record. Full runbook: docs/lab/citation-fidelity-v0.1.md.

Reproducibility
Question bank v0.1 (Q01–Q12), run date, engine version, prompt, and settings are recorded so a third party can repeat the run. Consumer panel (13 engines): Perplexity, Scira, Kagi, NotebookLM, Gemini, ChatGPT, Claude, Microsoft Copilot, DeepSeek, Le Chat, Consensus, Elicit, Manus. API-only tools documented separately.
Receipts

Sources and further reading

Student workflow

Citation Fidelity is the Lab protocol for checking whether cited claims match their sources. Pair it with the trusted-source brief before you submit.

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.