The rules that earn a verdict.
Quality gate
- Flagship content must include caveats and source freshness.
- Benchmark claims require task versions, prompts, failures, and safe raw examples.
- Pricing and privacy notes must show checked dates before public recommendation.
- Every claim in a dossier maps to a verifiable source.
Commercial guardrails
- Affiliate links are disclosed when present and never used on benchmark evidence rows.
- Sponsors cannot influence verdicts.
- Safety, privacy, and methodology remain public.
- Editorial corrections are logged, not silently rewritten.
How the 0–100 lands
The headline editorial fit score is a weighted roll-up of fourteen dossier dimensions — not a star rating and not a vendor leaderboard. Each dimension is 0–100 with a short note explaining usage ceiling (how much you can do before paying), cost-to-value (price vs rivals), opportunity cost (what you give up if you skip it), and competitive position (named alternatives at a glance).
Weights: wedge task fit 8% · feature depth 5% · free-tier utility 9% · cost-to-value 9% · opportunity cost 8% · source grounding 9% · privacy posture 9% · failure transparency 8% · evidence strength 8% · integration reach 6% · setup friction 6% · reliability 6% · competitive position 9%. The editorial-fit metric equals the headline.
Worked example — ChatGPT (synthesized evidence, public pricing): Free tier exists but caps power-user research → free-tier utility in the 60s. Pro at ~$20/mo vs Gemini and Claude → cost-to-value in the 70s. Cloud-only with export limits → opportunity cost moderated. Strong general wedge but not a citation-first research front door → wedge-task fit high, source grounding moderate. Dimensions roll up to a headline in the low 70s; solid tier until a full desk pass upgrades evidence.
Cards marked editorial · hands-on or self-run may override public-fact baselines after a structured trial. Rescored cards carry a verification stamp until the desk signs off.
Card stat bars
The six bars on each trading card (privacy, value, source quality, etc.) are deterministic heuristics derived from dossier text, not live benchmarks. The 0–100 editorial fit score is separate and hand-set. When self-run benchmark evidence exists, it appears in the dossier body, not as a hidden multiplier on the bars.
Compare verdicts
Head-to-head pages assemble the same public fields both cards ship with. The closing verdict is editorial logic in code, always traceable to pricing, privacy facets, capabilities, and failure modes you can read on each dossier.
VerdictPal Lab
First-party benchmarks run on the desk — with a frozen protocol, question bank, and runbook before any score ships. The Lab is not a vendor leaderboard; it tests the failure modes that actually burn students and researchers.
Evidence levels
- Self-run: VerdictPal ran the task. Logs and examples published.
- Editorial · hands-on: Editor used the tool, recorded observations.
- Synthesized: Compiled from multiple independent reports.
- Vendor claim: Reported but unverified. Treated with caution.
Status labels
- Flagship: Passes the full gate. Carries an editorial verdict.
- Solid: Reliable for the documented use case, with caveats.
- Lightweight: Editorial sketch. Needs deeper evidence to advance.
- Draft: Public scaffolding. Not a recommendation yet.
The Pack · Editorial newsletter
New cards in your inbox. Free.
One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.