VerdictPal · editorial desk · 2026VerdictPal

The rules that earn a verdict.

VerdictPal separates hands-on impressions, synthesized evidence, vendor claims, and self-run benchmark results. Every dossier wears its evidence label in public. For who runs the desk and what we refuse to do, see our trust page.

Quality gate

  • Flagship content must include caveats and source freshness.
  • Benchmark claims require task versions, prompts, failures, and safe raw examples.
  • Pricing and privacy notes must show checked dates before public recommendation.
  • Every claim in a dossier maps to a verifiable source.

Commercial guardrails

  • Affiliate links are disclosed when present and never used on benchmark evidence rows.
  • Sponsors cannot influence verdicts.
  • Safety, privacy, and methodology remain public.
  • Editorial corrections are logged, not silently rewritten.
Editorial principle: We would rather publish a small atlas of trustworthy dossiers than a large ranking we cannot defend. Empty is fine. Confident-and-wrong is not.

How the 0–100 lands

The headline editorial fit score is a weighted roll-up of fourteen dossier dimensions — not a star rating and not a vendor leaderboard. Each dimension is 0–100 with a short note explaining usage ceiling (how much you can do before paying), cost-to-value (price vs rivals), opportunity cost (what you give up if you skip it), and competitive position (named alternatives at a glance).

Weights: wedge task fit 8% · feature depth 5% · free-tier utility 9% · cost-to-value 9% · opportunity cost 8% · source grounding 9% · privacy posture 9% · failure transparency 8% · evidence strength 8% · integration reach 6% · setup friction 6% · reliability 6% · competitive position 9%. The editorial-fit metric equals the headline.

Worked example — ChatGPT (synthesized evidence, public pricing): Free tier exists but caps power-user research → free-tier utility in the 60s. Pro at ~$20/mo vs Gemini and Claude → cost-to-value in the 70s. Cloud-only with export limits → opportunity cost moderated. Strong general wedge but not a citation-first research front door → wedge-task fit high, source grounding moderate. Dimensions roll up to a headline in the low 70s; solid tier until a full desk pass upgrades evidence.

Cards marked editorial · hands-on or self-run may override public-fact baselines after a structured trial. Rescored cards carry a verification stamp until the desk signs off.

Card stat bars

The six bars on each trading card (privacy, value, source quality, etc.) are deterministic heuristics derived from dossier text, not live benchmarks. The 0–100 editorial fit score is separate and hand-set. When self-run benchmark evidence exists, it appears in the dossier body, not as a hidden multiplier on the bars.

Compare verdicts

Head-to-head pages assemble the same public fields both cards ship with. The closing verdict is editorial logic in code, always traceable to pricing, privacy facets, capabilities, and failure modes you can read on each dossier.

VerdictPal Lab

First-party benchmarks run on the desk — with a frozen protocol, question bank, and runbook before any score ships. The Lab is not a vendor leaderboard; it tests the failure modes that actually burn students and researchers.

Open the Lab · Citation Fidelity v0.1

Evidence levels

  • Self-run: VerdictPal ran the task. Logs and examples published.
  • Editorial · hands-on: Editor used the tool, recorded observations.
  • Synthesized: Compiled from multiple independent reports.
  • Vendor claim: Reported but unverified. Treated with caution.

Status labels

  • Flagship: Passes the full gate. Carries an editorial verdict.
  • Solid: Reliable for the documented use case, with caveats.
  • Lightweight: Editorial sketch. Needs deeper evidence to advance.
  • Draft: Public scaffolding. Not a recommendation yet.

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.