Why this benchmark is useful
It gives a compact signal for a specific capability. Use it as one dated receipt beside pricing, privacy, and hands-on evidence.
Benchmarks / Knowledge
When an answer-engine cites a source, does the source actually say what the answer claims? We grade real research questions on whether every cited claim is supported by the linked page — the failure mode students get burned by most.
It gives a compact signal for a specific capability. Use it as one dated receipt beside pricing, privacy, and hands-on evidence.
Read Citation Fidelity (VerdictPal Lab) as a active signal with low contamination risk. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.
No results published yet.
The protocol is public before the run ships.
Protocol v0.1 frozen. July 2026 desk run: 156/156 capture slots filled across 12 questions and 13 engines — human grading and sign-off pending before any score ships on the atlas.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
Frozen protocol v0.1 (2026-07-06). Twelve questions (Q01–Q12): three policy brief, three product comparison, three academic-adjacent, three local/current. Each engine gets the same prompt and default settings; two reviewers grade every cited claim independently (supported / unsupported / broken-link); disagreements are resolved on the record. Full runbook: docs/lab/citation-fidelity-v0.1.md.
Citation Fidelity is the Lab protocol for checking whether cited claims match their sources. Pair it with the trusted-source brief before you submit.
The Pack · Editorial newsletter
One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.