Benchmarks / Knowledge

Citation Fidelity (VerdictPal Lab)

When an answer-engine cites a source, does the source actually say what the answer claims? We grade real research questions on whether every cited claim is supported by the linked page — the failure mode students get burned by most.

KnowledgeDraftActiveLow contaminationVerdictPal LabSince 2026
What this does not measure
  • Answer quality or fluency — a beautifully written answer with a fabricated citation fails here on purpose.
  • Coverage — we test whether what is cited is true, not whether the model found every relevant source.
  • It is a small, hand-graded set — treat early numbers as directional, not definitive.
Analysis

Why this benchmark is useful

It gives a compact signal for a specific capability. Use it as one dated receipt beside pricing, privacy, and hands-on evidence.

Scope

Coverage map

Task family
Knowledge
Format
A fixed bank of 12 real student/researcher questions (policy brief, product comparison, academic-adjacent, local/current), run through each answer-engine; every cited claim is checked against its linked source by two reviewers.
Scoring
Citation-supported rate = supported cited claims / total cited claims, plus counts for unsupported and broken-link claims.
Maintainer
VerdictPal desk
Reading guide

How to read the scores

Read Citation Fidelity (VerdictPal Lab) as a active signal with low contamination. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.

Blind spots

What it does not cover

  • Answer quality or fluency — a beautifully written answer with a fabricated citation fails here on purpose.
  • Coverage — we test whether what is cited is true, not whether the model found every relevant source.
  • It is a small, hand-graded set — treat early numbers as directional, not definitive.
Scores

Evidence ledger

0 rows

No results published yet.

The protocol is public before the run ships.

Method

What it covers

Protocol v0.1 frozen. July 2026 desk run: 156/156 capture slots filled across 12 questions and 13 engines — human grading and sign-off pending before any score ships on the atlas.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

How we run it

Frozen protocol v0.1 (2026-07-06). Twelve questions (Q01–Q12): three policy brief, three product comparison, three academic-adjacent, three local/current. Each engine gets the same prompt and default settings; two reviewers grade every cited claim independently (supported / unsupported / broken-link); disagreements are resolved on the record. Full runbook: docs/lab/citation-fidelity-v0.1.md.

Reproducibility
Question bank v0.1 (Q01–Q12), run date, engine version, prompt, and settings are recorded so a third party can repeat the run. Consumer panel (13 engines): Perplexity, Scira, Kagi, NotebookLM, Gemini, ChatGPT, Claude, Microsoft Copilot, DeepSeek, Le Chat, Consensus, Elicit, Manus. API-only tools documented separately.
Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does Citation Fidelity (VerdictPal Lab) measure?

When an answer-engine cites a source, does the source actually say what the answer claims? We grade real research questions on whether every cited claim is supported by the linked page — the failure mode students get burned by most.

What does a high Citation Fidelity (VerdictPal Lab) score not prove?

A strong Citation Fidelity (VerdictPal Lab) result says nothing about:

  • Answer quality or fluency — a beautifully written answer with a fabricated citation fails here on purpose.
  • Coverage — we test whether what is cited is true, not whether the model found every relevant source.
  • It is a small, hand-graded set — treat early numbers as directional, not definitive.

How is Citation Fidelity (VerdictPal Lab) scored?

Citation-supported rate = supported cited claims / total cited claims, plus counts for unsupported and broken-link claims. Task format: A fixed bank of 12 real student/researcher questions (policy brief, product comparison, academic-adjacent, local/current), run through each answer-engine; every cited claim is checked against its linked source by two reviewers.

Is Citation Fidelity (VerdictPal Lab) saturated?

Citation Fidelity (VerdictPal Lab) is currently marked Active in the atlas.

Can Citation Fidelity (VerdictPal Lab) results be contaminated by training data?

Contamination risk for Citation Fidelity (VerdictPal Lab) is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

Student workflow

Citation Fidelity is the Lab protocol for checking whether cited claims match their sources. Pair it with the trusted-source brief before you submit.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.