VerdictPal · editorial desk · 2026VerdictPal
Freshness

What changed

Card changelog entries and Pack drops, grouped by week.

2026-W29

  1. GPT-5.6 Luna

    Recorded general availability across ChatGPT, Codex, and the API and refreshed benchmark evidence.

    Model
  2. GPT-5.6 Sol

    Recorded general availability across ChatGPT, Codex, and the API, then refreshed Vals and Artificial Analysis benchmark evidence.

    Model
  3. GPT-5.6 Terra

    Recorded general availability across ChatGPT, Codex, and the API and refreshed benchmark evidence.

    Model

2026-W28

  1. AI Builder Design Test on the Lab bench — four builders, same prompts, very different results

    VerdictPal Lab's second first-party protocol: four identical multi-turn design prompts through Lovable, v0, Replit, and Bolt. Run 01 recorded a Turn 1 ranking (Lovable → v0 → Replit → Bolt) and the cost-of-arrival in each vendor's native unit — no padded leaderboard, formal scored rows pending a repeatable rubric. /benchmarks/ai-builder-design-test · /benchmarks/lab

    Drop
  2. Grok 4.5

    Grok 4.5 appears on Artificial Analysis (high-effort index 53.8).

    Model
  3. Citation Fidelity v0.1 protocol frozen on the Lab bench

    VerdictPal Lab's first first-party benchmark: twelve research questions, dual-reviewer citation grading, frozen runbook on the atlas card — protocol published, desk run still pending, no padded leaderboard. /benchmarks/citation-fidelity · /benchmarks/lab

    Drop
  4. Claude

    Evidence integrity fix: reverted false editorial-hands-on upgrade; evidence stays synthesized and qualityGate downgraded flagship → solid per docs/EDITORIAL-SOP.md §2 evidence ladder. No desk hands-on trial on file; benchmarkReady remains false.

    Card

2026-W27

  1. Five research guides now on the atlas

    Citation audit, systematic review protocol, data cleaning pipeline, grant writer path, and peer reviewer path — solid guides with step cards and verification checklists.

    Drop
  2. GPT-5.6 Luna

    Updated public pricing: Luna is listed at $1 input / $6 output per 1M tokens.

    Model
  3. GPT-5.6 Sol

    Updated public pricing: Sol is listed at $5 input / $30 output per 1M tokens, with 1.25x cache-write billing and 90% cached-read discount.

    Model
  4. GPT-5.6 Terra

    Updated public pricing: Terra is listed at $2.50 input / $15 output per 1M tokens.

    Model
  5. OpenRouter

    Added LongCat-2.0/Owl Alpha release context: Meituan identified the OpenRouter stealth route as LongCat-backed after it had already become a high-volume coding model. Treat it as evidence of routing-market adoption, not as a privacy shortcut or independently verified benchmark row.

    Card
  6. Claude Sonnet 5

    Claude Sonnet 5 listed on Artificial Analysis (max-effort index 53.4).

    Model
  7. LongCat-2.0

    Meituan revealed LongCat-2.0 and identified OpenRouter's Owl Alpha stealth route as LongCat-backed; GitHub/Hugging Face pages say model weights are coming soon.

    Model

2026-W26

  1. GPT-5.6 Luna

    Previewed as the fastest and most cost-efficient GPT-5.6 tier.

    Model
  2. GPT-5.6 Sol

    OpenAI previewed GPT-5.6 Sol alongside Terra and Luna.

    Model
  3. GPT-5.6 Terra

    Previewed as the balanced GPT-5.6 tier.

    Model
  4. Student research workflow on the guides pillar

    Path, stack, playbook, and showcase for turning broad prompts into source-backed briefs — the workflow we point new readers at first.

    Drop
  5. Factory flagship dossier shipped

    Agent-native delivery with Droid, Spec Mode, missions, and enterprise controls — fit 74 with the autonomy and pricing caveats on the record.

    Drop
  6. Cursor flagship dossier shipped

    VS Code-compatible AI editor at fit 81 — Tab, Chat, Agent, Privacy Mode, team governance, and the diff-review failure modes we publish before you merge.

    Drop
  7. Gemini flagship dossier shipped

    Google's multimodal assistant ecosystem at fit 75 — Search grounding, Deep Research, storage bundles, and where it is not a cited search engine.

    Drop

2026-W25

  1. NotebookLM flagship dossier shipped

    Source-grounded reading workspace at fit 78 — corpus chat, Audio Overviews, Flashcards, and the empty-corpus trap documented on the card.

    Drop
  2. Perplexity flagship dossier shipped

    Cited search front door at fit 72 — consumer Pro/Max, Deep Research, Sonar APIs, and failure modes for bibliographies you have not opened.

    Drop
  3. Cursor

    Production-readiness audit pass: editorsNote and pricing/privacy checked dates confirmed. benchmarkReady remains false intentionally until agent-diff safety, usage-burn, and Privacy Mode runs land.

    Card
  4. Factory

    Production-readiness audit pass: editorsNote and pricing/privacy checked dates confirmed. benchmarkReady remains false intentionally until Spec Mode, Droid Exec autonomy, and enterprise-control runs land.

    Card
  5. Gemini

    Production-readiness audit pass: editorsNote and pricing/privacy checked dates confirmed. No editorial claims changed; bench rows remain explicitly pending until the desk protocol runs (Next action in qualityGate.notes).

    Card
  6. Perplexity

    Production-readiness audit pass: editorsNote, pricing/privacy checked dates confirmed; Sonar/Search API token + per-request fee rows aligned with docs.perplexity.ai. benchmarkReady remains false intentionally until desk citation audit and Sonar latency runs land.

    Card
  7. Command Code

    Shipped Command Code flagship card — terminal agent with taste learning, open-model deals, and full pricing/privacy research from commandcode.ai.

    Card
  8. Command Code flagship dossier shipped

    Terminal coding agent with taste-1 learning, open-model deals from $1/mo, and the full pricing and privacy picture — including what lives in .commandcode/.

    Drop
  9. OpenCode

    Shipped OpenCode flagship card — OSS agent, Go/Zen pricing, Plan/Build modes, and Zen privacy exceptions documented.

    Card
  10. OpenCode flagship dossier shipped

    The open-source coding agent at fit 95 — Plan/Build modes, Go and Zen pricing, 75+ providers, and the free-model privacy caveats on the record.

    Drop
  11. Raycast

    Shipped Raycast as a flagship Knowledge work card — the macOS launcher we would install on every research Mac before debating another AI tab.

    Card
  12. Raycast flagship dossier shipped

    The macOS launcher we install before debating another AI tab — extensions, Quick AI, clipboard, and honest failure modes at fit 97.

    Drop

2026-W24

  1. GLM-5.2

    Released GLM-5.2 with 1M context, 128K max output, two thinking modes, and long-horizon coding focus. MIT open weights announced for the following week.

    Model
  2. Kimi K2.7 Code

    Released Kimi K2.7 Code with 256K context, mandatory thinking mode, multimodal tool examples, and reported 30% lower reasoning-token usage versus K2.6.

    Model
  3. Canva Business

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  4. ChatGPT

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  5. ChatPRD

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  6. Claude Fable 5

    Released Claude Fable 5 — first GA Mythos-class model; Vals Index 75.15%.

    Model
  7. Connected Papers

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  8. Consensus

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  9. Crossref

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  10. DeepSeek

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  11. ElevenLabs

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  12. Elicit

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  13. Exa

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  14. Firecrawl

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  15. Gamma

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  16. Google Scholar

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  17. Grammarly

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  18. Gumloop

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  19. Kagi

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  20. Linear

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  21. Litmaps

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  22. Lovable

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  23. Magic Patterns

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  24. Manus

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  25. Mendeley

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  26. Microsoft Copilot

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  27. Mistral Le Chat

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  28. Mobbin

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  29. n8n

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  30. Obsidian

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  31. OpenAlex

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  32. OpenRouter

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  33. Overleaf

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  34. Paperpile

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  35. PostHog

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  36. Rayyan

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  37. Replit

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  38. Research Rabbit

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  39. Scholarcy

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  40. Scira AI

    Refreshed pricing evidence from scira.ai/about and api.scira.ai/pricing, added machine-readable pricing tiers, and marked benchmark evidence ready for the public dossier while keeping individual desk benchmark rows explicit about what still needs live runs.

    Card
  41. Scite

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  42. Semantic Scholar

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  43. Warp

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  44. Wispr Flow

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  45. You.com

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  46. Zotero

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Card
  47. Grant deadline set shipped

    A flagship path, stack, playbook, and walkthrough for moving from a blank Specific Aims page to a submission-ready one-pager and budget justification in under three hours.

    Drop

2026-W23

  1. ChatGPT

    ChatGPT is the baseline general assistant card, but it now needs to be framed as a family of surfaces: consumer ChatGPT, Business/Enterprise workspaces, Codex/agent features, and the separate OpenAI API. Its strength is breadth — drafting, file work, multimodal analysis, coding, agents, and extensions — while its risk is that fluent output can make weak sourcing, privacy assumptions, or plan limits invisible.

    Card
  2. Claude

    Claude is the careful-drafting and long-context card. The updated evidence reinforces its core lane: nuanced rewriting, document analysis, Projects, Artifacts, Research, Claude Code, Cowork, web search, and organization search/connectors. The editorial risk is quota and surface confusion: consumer Pro/Max, Team, Enterprise, and API/tool pricing are meaningfully different products.

    Card
  3. Connected Papers

    Connected Papers remains the clean visual literature-map card. It is especially useful for orienting around one seed paper, but its core limitation is still methodological: a graph is not a reproducible search strategy. This pass adds official plan/feature detail, group/library plans, payment processing, scholarship support, and free-quota boundaries.

    Card
  4. Consensus

    Consensus remains the claim-shaped academic Q&A card. It is strongest when the user has a research question and needs cited orientation quickly. This pass adds current corpus/user-scale claims, Research Agent, account/privacy controls, institutional full-text linking, and a sharper “answer box is not a literature review” warning.

    Card
  5. Crossref

    Crossref is scholarly infrastructure, not a shiny discovery app. It remains the DOI metadata verification card: use it for DOI records, member/publisher metadata, references, funder/license fields, relation metadata, and citation plumbing. This pass adds the 2025/2026 REST API rate-limit shift and the practical “polite pool” guidance that makes Crossref usable in responsible pipelines.

    Card
  6. Cursor

    Cursor remains the AI-native code editor flagship, but the card now needs to emphasize three separations: editor workflow versus headless agents, individual usage pools versus team/enterprise governance, and Privacy Mode versus ordinary provider transit. Cursor is excellent for repo-local Tab, Chat, Agent, MCP, cloud agents, and Bugbot workflows, but only if diffs are reviewed and tests run.

    Card
  7. DeepSeek

    DeepSeek is the low-cost reasoning/API card. This pass sharpens the key facts: DeepSeek’s API is OpenAI- and Anthropic-compatible, current V4-Flash/V4-Pro pricing is extremely low, context length is listed at 1M, and legacy deepseek-chat / deepseek-reasoner names are being deprecated. The low price is real, but the editorial warning is bigger than usual: privacy, residency, institutional policy, and geopolitical availability need explicit review.

    Card
  8. ElevenLabs

    ElevenLabs is the voice and audio AI card. This pass uses primary ElevenLabs pricing and privacy pages plus compliance/search results. The product’s quality can make outputs feel finished, but the editorial review must center voice consent, identity risk, disclosure, credit burn, and whether uploaded voices/interviews belong in a cloud voice system.

    Card
  9. Elicit

    Elicit is no longer just a paper-search assistant; it is positioning itself as an evidence-synthesis system for reports, systematic reviews, dynamic screening, extraction, and sentence-level citations over a very large scientific corpus. The richer data supports a higher-confidence card, but the recommendation still depends on benchmarked extraction accuracy against a hand-coded table.

    Card
  10. Exa

    Exa is one of the cleanest “search as infrastructure” cards in the VerdictPal stack. The current pricing and docs make the value clearer: real-time search, webpage text/highlights, configurable latency, contents retrieval, Answer endpoint, Deep Search, Monitors, and Agent runs. The core risk is economic and privacy-boundary creep when agents fan out across many searches, contents calls, summaries, and enrichment steps.

    Card
  11. Factory

    Factory is the agent-native software-delivery card. It is strongest when the task is bigger than autocomplete: repo understanding, spec-planned refactors, terminal/IDE delegation, background agents, SDK use, and CI-style Droid Exec workflows. This pass adds current 2026 Factory/Droid source detail and sharpens the enterprise/security boundary.

    Card
  12. Firecrawl

    Firecrawl is the web-extraction infrastructure card. It should be described as context extraction for agents and research pipelines, not as a generic search engine. This pass adds current credit economics, enterprise ZDR/SOC2 claims, endpoint credit costs, and the important warning that hosted scraping is always a trust, permission, and site-terms problem.

    Card
  13. Gamma

    Gamma is the AI presentation and visual-document first-draft card. This pass uses official Gamma pricing/product/privacy/search snippets, but the pricing and privacy pages timed out on full load, so the dossier clearly marks what is verified from retrieved official snippets and what needs final live-page verification before publication.

    Card
  14. Gemini

    Deepened from official Google pages: pricing across consumer (AI Plus/Pro/Ultra), Workspace add-on, Gemini API paid tier, and Vertex AI; privacy language for consumer vs API vs Workspace/Vertex; capability list refreshed to current Gemini app + API surface.

    Card
  15. Google Scholar

    Google Scholar remains the academic-search baseline: broad, familiar, fast, and useful for finding known items, PDFs, library links, citation trails, alerts, and case law. Its weakness is methodological opacity. It should be recommended as a discovery and citation-chasing habit, not as a reproducible search strategy by itself.

    Card
  16. Grammarly

    Grammarly is the revision-pass card, not the argument-writing card. This pass updates the card for Grammarly Pro and Superhuman-suite positioning: Free gives core correction plus 100 AI prompts, Pro adds rewrites/tone/brand features and 2,000 AI prompts, team features run up to 149 seats, and larger organizations move to Superhuman Go/enterprise-style controls. The core warning remains voice flattening and cloud text processing.

    Card
  17. Gumloop

    Gumloop is the no-code AI workflow and agent-automation card. This pass is based on primary Gumloop pricing and privacy pages, plus official product/search results. The review should focus on whether credit-based agent workflows stay understandable after real data, credentials, retries, and handoffs enter the loop.

    Card
  18. Kagi

    Kagi is the paid-search card: a user-funded, ad-free search engine with serious privacy engineering and enough Assistant functionality to be useful without turning into an AI slop layer. The strongest angle is incentive alignment — no ads, no result-click tracking, no analytics/telemetry on the main site, and paid plans that make the business model legible. The recommendation still depends on a 25-query search-quality benchmark against Google, Brave, DuckDuckGo, Perplexity, and You.com.

    Card
  19. Litmaps

    Litmaps is the citation-map plus monitoring card. It is valuable when a user starts from seed papers and needs to discover connected work, visualize the field, sync references, and keep watch for new papers. This pass fixes the malformed sources area and adds current Litmaps feature, coverage, and DPA details.

    Card
  20. Mendeley

    Mendeley is the Elsevier-backed cloud-sync reference manager with increasingly AI-shaped library features. The richer data clarifies the buyer’s question: Mendeley is capable for PDF import, annotation, Word citation, groups, and AI over a library, but the deciding tradeoff is Elsevier cloud trust, storage/AI plan economics, and future switching cost versus Zotero and Paperpile.

    Card
  21. Microsoft Copilot

    Microsoft Copilot is a family-name card, not a single-product card. This pass uses Microsoft Learn pages for Microsoft 365 Copilot and Copilot Chat data protection plus Microsoft pricing search results. The key editorial job is to separate consumer Copilot, Microsoft 365 Copilot Chat, paid Microsoft 365 Copilot, and Azure/OpenAI developer paths.

    Card
  22. Mistral Le Chat

    Mistral Le Chat is the EU-headquartered assistant comparison card. This pass strengthens the privacy and API boundary: Le Chat inputs/outputs are retained until account/conversation deletion, API inputs/outputs are generally kept for 30 rolling days for abuse monitoring unless zero data retention is activated, and Agents/Fine-tuning have separate retention rules. That makes Mistral useful to compare against US-default assistants, but not automatically compliant.

    Card
  23. n8n

    n8n is the research-ops automation card. It is strongest when a team needs visible workflow glue across APIs, webhooks, databases, AI steps, Notion, and alerts. The updated evidence makes the trade-off clearer: every plan now leans into unlimited users/workflows/steps, but pricing is based on full workflow executions, and self-hosted still means the operator owns patching, backups, secrets, telemetry choices, and failure recovery.

    Card
  24. NotebookLM

    NotebookLM remains the source-grounded workspace flagship. The strongest VerdictPal framing is not “AI search,” but “corpus-first reading workspace”: upload or discover sources, ask grounded questions, generate study artifacts, and verify claims against citations. This pass adds current Enterprise, Workspace, and Google AI plan distinctions so the privacy story is clearer.

    Card
  25. Obsidian

    Obsidian is the local-first knowledge-base card. The important update is that Obsidian is now free for work as well as personal use, while paid services remain optional add-ons for Sync and Publish. That strengthens the card’s editorial angle: Obsidian is not a cloud workspace trying to lock your notes away; it is a Markdown vault where the user owns the files, but also owns the backup, plugin, and structure decisions.

    Card
  26. OpenAlex

    OpenAlex is the open scholarly graph card. It sits between Crossref and Semantic Scholar: broader discovery and mapping than DOI registry metadata, more reproducible and open than closed search products, but still an aggregated graph that requires verification for final citations. This pass adds the 2026 usage-based pricing and privacy promise details.

    Card
  27. OpenRouter

    OpenRouter is the model-router card. It is useful when a developer wants one OpenAI-compatible API surface for many LLMs, fallback routing, model comparisons, BYOK, and spend controls. The updated card sharpens the key caveat: routing convenience adds a platform fee and a subprocessor/provider-retention surface that must be controlled before confidential research or user data flows through it.

    Card
  28. Overleaf

    Overleaf is the collaborative LaTeX card. This pass updates the card with current plan evidence, including AI allowance language, collaborator limits, 24x compile timeout on paid tiers, real-time track changes, and the subscriber-owned project entitlement model. The core trade remains: Overleaf removes local TeX and collaboration friction, but your project source and compile workflow live in a cloud editor.

    Card
  29. Paperpile

    Paperpile is the convenience-first reference manager for Google Docs and browser-native writing. The stronger data now makes the tradeoff clearer: it has excellent Google Docs and PDF workflow coverage, annual pricing only, unlimited PDF storage on paid plans, and a cloud/Google-account privacy boundary that should be explicit for thesis and institutional work.

    Card
  30. Perplexity

    Perplexity remains the cited-answer benchmark for consumer AI search, but the card now needs to treat Perplexity as two related products: the consumer answer engine and the API platform. The consumer product is useful for fast source maps and Deep Research drafts; the API platform is a separate developer surface with Sonar, Search, Agent, and Embeddings APIs. The hard editorial rule stays the same: citations are leads, not bibliography-ready evidence.

    Card
  31. PostHog

    PostHog is the product-analytics and experimentation-stack card. It belongs in VerdictPal because product analytics, session replay, feature flags, experiments, surveys, pipelines, and AI/product ops determine how the product learns after launch. The recommendation stays gated because PostHog’s biggest risks are exactly the ones product teams underestimate: event volume, replay privacy, retention, and usage-based cost.

    Card
  32. Rayyan

    Rayyan is a focused systematic-review screening workspace with stronger plan mechanics and AI boundaries than the earlier card captured. The product’s editorial value is title/abstract screening, deduplication, PICO extraction, PRISMA support, reviewer/viewer collaboration, mobile work, and institutional ResearchPilot. The gold-standard warning is unchanged: AI can assist review logistics, but the final inclusion/exclusion decision must remain human and auditable.

    Card
  33. Replit

    Replit is the browser-based AI build-and-deploy card. It is useful when a user wants one place for ideation, agent building, database, hosting, collaboration, and publishing. The review must focus on code quality, credit burn, deployment security, and whether the user understands what the agent built.

    Card
  34. Research Rabbit

    Research Rabbit is the visual exploration and “follow the trail” card. It belongs beside Connected Papers and Litmaps, but its emphasis is adaptive exploration, collections, and seeing how papers/authors/concepts connect. This pass adds current feature language, data-source hints, and a clearer privacy/method boundary.

    Card
  35. Scholarcy

    Scholarcy is the paper-skimming and structured-summary card. It can save time for triage, accessibility, and organization, but it must be reviewed as a reading aid, not a reading replacement. This pass adds current product/pricing/help details, browser-extension behavior, export/library capabilities, and a sharper copyright/privacy boundary for uploaded PDFs.

    Card
  36. Scira AI

    Scira remains the VerdictPal flagship template because it combines a cited search UI, open-source AGPL codebase, hosted Free/Pro/Max plans, an API platform, MCP surface, and a deep dossier structure. This pass replaces stale/malformed page content with a cleaner benchmarkable dossier while preserving the core stance: strong public recommendation, but benchmark-ready still false until the pending desk rows are actually run.

    Card
  37. Scite

    Scite is the citation-context card. It is valuable because it asks a better question than raw citation count: did later papers support, contrast, or merely mention the claim? This pass strengthens the card with Scite’s current Smart Citations, MCP, API, publisher/full-text coverage, privacy-policy, and reference-check caveats. The public recommendation remains gated until label accuracy is benchmarked against actual citing passages.

    Card
  38. Semantic Scholar

    Semantic Scholar remains a flagship academic-search layer because it combines a free scholarly search UI with TLDRs, author pages, alerts, a developer API, and downloadable scholarly graph data. It is more structured than Google Scholar and more paper-discovery oriented than Crossref. The warning is unchanged: AI summaries and graph edges are discovery aids, not evidence.

    Card
  39. Warp

    Warp is the agentic terminal/workspace card. It is compelling because it puts local and cloud coding agents where developers already run commands, but the evaluation must center command safety, repo hygiene, and data controls.

    Card
  40. You.com

    You.com should now be reviewed primarily as a web-search API platform for AI builders, not as a nostalgic consumer search/chat competitor. Its current public pitch centers on Web Search APIs, Research API, zero-data-retention options, SOC2, DPA readiness, high rate limits, and enterprise deployment. The card’s benchmark needs to test API source quality, freshness, cost, and data-control clarity against Exa, Perplexity Sonar/Search, Brave, Tavily-style APIs, and Kagi.

    Card
  41. Zotero

    Zotero remains the strongest local-first default for student and academic reference management. The new data points reinforce the core positioning: the desktop workflow works without an account, sync is optional and disabled by default, paid storage funds a nonprofit project, and the main editorial risk is not capability but governance around sync, attachments, plugins, and group ownership.

    Card
  42. Canva Business

    Imported from the Notion Card pipeline with Scira-gate status corrected to solid.

    Card
  43. ChatPRD

    Imported Notion state-of-art research and lowered the editorial fit to reflect weak desk usefulness plus Teams-gated Linear integration.

    Card
  44. Linear

    Imported from the Notion Card pipeline with Scira-gate status corrected to solid.

    Card
  45. Lovable

    Imported from the Notion Card pipeline with Scira-gate status corrected to solid.

    Card
  46. Magic Patterns

    Imported from the Notion Card pipeline with Scira-gate status corrected to solid.

    Card
  47. Manus

    Imported Notion state-of-art fields and corrected Scira-gate status to solid.

    Card
  48. Mobbin

    Imported from the Notion Card pipeline with Scira-gate status corrected to solid.

    Card
  49. Wispr Flow

    Imported from the Notion Card pipeline with Scira-gate status corrected to solid.

    Card
  50. Dark mode rebuilt for the evidence instrument

    Dark surfaces now use muted amber, readable borders, and component-specific overrides instead of neon cards and inverted cream panels.

    Drop
  51. Trusted-source brief workflow shipped

    The Guides pillar now has a real student research path, tool stack, playbook, and showcase for turning broad prompts into source-backed briefs.

    Drop
  52. Claude Mythos Preview

    Released Claude Mythos research preview — tops Vals (73.42%) and SWE-bench Verified (93.9%, best-of-3).

    Model

2026-W22

  1. MiniMax M3

    Released MiniMax M3 — Terminal-Bench 2.0 66.0%, 1M context.

    Model
  2. Claude Opus 4.8

    Released Claude Opus 4.8 — top of Vals (70.17%) and AA (61.4).

    Model

2026-W21

  1. Qwen 3.7 Max

    Released Qwen 3.7 Max — AA Index 56.6, Terminal-Bench 2.0 69.7%.

    Model
  2. Gemini 3.5 Flash

    Released Gemini 3.5 Flash — MCP Atlas leader (83.6%), 1M context, cheapest frontier pricing.

    Model

2026-W18

  1. Grok 4.3

    Released Grok 4.3 with 2M context and real-time X search.

    Model

2026-W17

  1. DeepSeek V4 Pro Max

    Released DeepSeek V4 Pro Max with open weights and frontier coding scores.

    Model
  2. GPT 5.5

    Released GPT 5.5 with native audio, image, and 512K context.

    Model
  3. GPT-5.5 Pro

    Released GPT-5.5 Pro — premium tier, $30/$180, Responses API.

    Model
  4. Kimi K2.6

    Released Kimi K2.6 — Vals #5 at 55.55%, GPQA Diamond 90.5% (open-source leader).

    Model

2026-W16

  1. Claude Opus 4.7

    Released Claude Opus 4.7 — 13% coding lift, 3x production tasks, 3.75 MP vision.

    Model
  2. Claude Opus 4.8

    Claude Opus 4.7 retired for general availability.

    Model

2026-W15

  1. GLM-5.1

    Released GLM-5.1 with 200K context, 128K max output, long-horizon coding focus, and MIT-licensed open weights.

    Model

2026-W14

  1. Qwen 3.7 Max

    Qwen 3.7 Plus retired for general availability.

    Model

2026-W12

  1. GPT-5.4 Mini

    Released GPT-5.4 Mini — best price/performance, $0.75/$4.50.

    Model
  2. GPT-5.4 Nano

    Released GPT-5.4 Nano — budget tier, $0.20/$1.25.

    Model

2026-W11

  1. MiniMax M3

    MiniMax M2.7 retired as the flagship open-weights model.

    Model

2026-W10

  1. GPT 5.4

    Released GPT 5.4 — AA Index 56.8, Terminal-Bench 2.0 81.8% (ForgeCode).

    Model
  2. GPT 5.5

    GPT 5.4 retired for general availability.

    Model
  3. GPT-5.4 Pro

    Released GPT-5.4 Pro — premium tier, $30/$180.

    Model
  4. Grok 4.3

    Grok 4.20 retired from general availability.

    Model

2026-W08

  1. Gemini 3.1 Pro

    Released Gemini 3.1 Pro — AA Index 57.2, ARC-AGI-2 77.1%.

    Model
  2. Claude Sonnet 4.6

    Released Claude Sonnet 4.6 — Vals #3 at 60.30%, 1M context, $3 / $15.

    Model

2026-W07

  1. DeepSeek V4 Pro Max

    DeepSeek V3.2 retired as the flagship open-weights model.

    Model

2026-W06

  1. Claude Opus 4.6

    Released Claude Opus 4.6 — first Opus with 1M context, 128K output, agent teams.

    Model

2025-W50

  1. GPT 5.4

    GPT 5.2 retired for general availability.

    Model
  2. GPT-5.2

    Released GPT-5.2 — 410K context, configurable reasoning, $1.75/$14.

    Model

2025-W48

  1. Claude Opus 4.5

    Released Claude Opus 4.5 — best coding and agent model at launch, $5/$25.

    Model

2025-W40

  1. Claude Haiku 4.5

    Released Claude Haiku 4.5 — speed tier, $1/$5, 200K context.

    Model
Pack drops feed