Why this benchmark is useful
It gives a compact signal for a specific capability. Use it as one dated receipt beside pricing, privacy, and hands-on evidence.
When an answer-engine cites a source, does the source actually say what the answer claims? We grade real research questions on whether every cited claim is supported by the linked page — the failure mode students get burned by most.
It gives a compact signal for a specific capability. Use it as one dated receipt beside pricing, privacy, and hands-on evidence.
Read Citation Fidelity (VerdictPal Lab) as a active signal with low contamination. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.
No results published yet.
The protocol is public before the run ships.
Protocol v0.1 frozen. July 2026 desk run: 156/156 capture slots filled across 12 questions and 13 engines — human grading and sign-off pending before any score ships on the atlas.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
Frozen protocol v0.1 (2026-07-06). Twelve questions (Q01–Q12): three policy brief, three product comparison, three academic-adjacent, three local/current. Each engine gets the same prompt and default settings; two reviewers grade every cited claim independently (supported / unsupported / broken-link); disagreements are resolved on the record. Full runbook: docs/lab/citation-fidelity-v0.1.md.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
When an answer-engine cites a source, does the source actually say what the answer claims? We grade real research questions on whether every cited claim is supported by the linked page — the failure mode students get burned by most.
A strong Citation Fidelity (VerdictPal Lab) result says nothing about:
Citation-supported rate = supported cited claims / total cited claims, plus counts for unsupported and broken-link claims. Task format: A fixed bank of 12 real student/researcher questions (policy brief, product comparison, academic-adjacent, local/current), run through each answer-engine; every cited claim is checked against its linked source by two reviewers.
Citation Fidelity (VerdictPal Lab) is currently marked Active in the atlas.
Contamination risk for Citation Fidelity (VerdictPal Lab) is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
Citation Fidelity is the Lab protocol for checking whether cited claims match their sources. Pair it with the trusted-source brief before you submit.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.