What changed
2026-W29
- GPT-5.6 Luna
Recorded general availability across ChatGPT, Codex, and the API and refreshed benchmark evidence.
Model - GPT-5.6 Sol
Recorded general availability across ChatGPT, Codex, and the API, then refreshed Vals and Artificial Analysis benchmark evidence.
Model - GPT-5.6 Terra
Recorded general availability across ChatGPT, Codex, and the API and refreshed benchmark evidence.
Model
2026-W28
- AI Builder Design Test on the Lab bench — four builders, same prompts, very different results
VerdictPal Lab's second first-party protocol: four identical multi-turn design prompts through Lovable, v0, Replit, and Bolt. Run 01 recorded a Turn 1 ranking (Lovable → v0 → Replit → Bolt) and the cost-of-arrival in each vendor's native unit — no padded leaderboard, formal scored rows pending a repeatable rubric. /benchmarks/ai-builder-design-test · /benchmarks/lab
Drop - Citation Fidelity v0.1 protocol frozen on the Lab bench
VerdictPal Lab's first first-party benchmark: twelve research questions, dual-reviewer citation grading, frozen runbook on the atlas card — protocol published, desk run still pending, no padded leaderboard. /benchmarks/citation-fidelity · /benchmarks/lab
Drop - Claude
Evidence integrity fix: reverted false editorial-hands-on upgrade; evidence stays synthesized and qualityGate downgraded flagship → solid per docs/EDITORIAL-SOP.md §2 evidence ladder. No desk hands-on trial on file; benchmarkReady remains false.
Card
2026-W27
- Five research guides now on the atlas
Citation audit, systematic review protocol, data cleaning pipeline, grant writer path, and peer reviewer path — solid guides with step cards and verification checklists.
Drop - GPT-5.6 Sol
Updated public pricing: Sol is listed at $5 input / $30 output per 1M tokens, with 1.25x cache-write billing and 90% cached-read discount.
Model - GPT-5.6 Terra
Updated public pricing: Terra is listed at $2.50 input / $15 output per 1M tokens.
Model - OpenRouter
Added LongCat-2.0/Owl Alpha release context: Meituan identified the OpenRouter stealth route as LongCat-backed after it had already become a high-volume coding model. Treat it as evidence of routing-market adoption, not as a privacy shortcut or independently verified benchmark row.
Card - LongCat-2.0
Meituan revealed LongCat-2.0 and identified OpenRouter's Owl Alpha stealth route as LongCat-backed; GitHub/Hugging Face pages say model weights are coming soon.
Model
2026-W26
- Student research workflow on the guides pillar
Path, stack, playbook, and showcase for turning broad prompts into source-backed briefs — the workflow we point new readers at first.
Drop - Factory flagship dossier shipped
Agent-native delivery with Droid, Spec Mode, missions, and enterprise controls — fit 74 with the autonomy and pricing caveats on the record.
Drop - Cursor flagship dossier shipped
VS Code-compatible AI editor at fit 81 — Tab, Chat, Agent, Privacy Mode, team governance, and the diff-review failure modes we publish before you merge.
Drop - Gemini flagship dossier shipped
Google's multimodal assistant ecosystem at fit 75 — Search grounding, Deep Research, storage bundles, and where it is not a cited search engine.
Drop
2026-W25
- NotebookLM flagship dossier shipped
Source-grounded reading workspace at fit 78 — corpus chat, Audio Overviews, Flashcards, and the empty-corpus trap documented on the card.
Drop - Perplexity flagship dossier shipped
Cited search front door at fit 72 — consumer Pro/Max, Deep Research, Sonar APIs, and failure modes for bibliographies you have not opened.
Drop - Cursor
Production-readiness audit pass: editorsNote and pricing/privacy checked dates confirmed. benchmarkReady remains false intentionally until agent-diff safety, usage-burn, and Privacy Mode runs land.
Card - Factory
Production-readiness audit pass: editorsNote and pricing/privacy checked dates confirmed. benchmarkReady remains false intentionally until Spec Mode, Droid Exec autonomy, and enterprise-control runs land.
Card - Gemini
Production-readiness audit pass: editorsNote and pricing/privacy checked dates confirmed. No editorial claims changed; bench rows remain explicitly pending until the desk protocol runs (Next action in qualityGate.notes).
Card - Perplexity
Production-readiness audit pass: editorsNote, pricing/privacy checked dates confirmed; Sonar/Search API token + per-request fee rows aligned with docs.perplexity.ai. benchmarkReady remains false intentionally until desk citation audit and Sonar latency runs land.
Card - Command Code
Shipped Command Code flagship card — terminal agent with taste learning, open-model deals, and full pricing/privacy research from commandcode.ai.
Card - Command Code flagship dossier shipped
Terminal coding agent with taste-1 learning, open-model deals from $1/mo, and the full pricing and privacy picture — including what lives in .commandcode/.
Drop - OpenCode
Shipped OpenCode flagship card — OSS agent, Go/Zen pricing, Plan/Build modes, and Zen privacy exceptions documented.
Card - OpenCode flagship dossier shipped
The open-source coding agent at fit 95 — Plan/Build modes, Go and Zen pricing, 75+ providers, and the free-model privacy caveats on the record.
Drop - Raycast
Shipped Raycast as a flagship Knowledge work card — the macOS launcher we would install on every research Mac before debating another AI tab.
Card - Raycast flagship dossier shipped
The macOS launcher we install before debating another AI tab — extensions, Quick AI, clipboard, and honest failure modes at fit 97.
Drop
2026-W24
- GLM-5.2
Released GLM-5.2 with 1M context, 128K max output, two thinking modes, and long-horizon coding focus. MIT open weights announced for the following week.
Model - Kimi K2.7 Code
Released Kimi K2.7 Code with 256K context, mandatory thinking mode, multimodal tool examples, and reported 30% lower reasoning-token usage versus K2.6.
Model - Canva Business
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - ChatGPT
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - ChatPRD
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Connected Papers
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Consensus
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Crossref
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - DeepSeek
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - ElevenLabs
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Elicit
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Exa
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Firecrawl
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Gamma
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Google Scholar
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Grammarly
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Gumloop
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Kagi
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Linear
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Litmaps
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Lovable
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Magic Patterns
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Manus
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Mendeley
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Microsoft Copilot
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Mistral Le Chat
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Mobbin
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - n8n
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Obsidian
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - OpenAlex
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - OpenRouter
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Overleaf
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Paperpile
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - PostHog
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Rayyan
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Replit
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Research Rabbit
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Scholarcy
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Scira AI
Refreshed pricing evidence from scira.ai/about and api.scira.ai/pricing, added machine-readable pricing tiers, and marked benchmark evidence ready for the public dossier while keeping individual desk benchmark rows explicit about what still needs live runs.
Card - Scite
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Semantic Scholar
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Warp
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Wispr Flow
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - You.com
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Zotero
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Card - Grant deadline set shipped
A flagship path, stack, playbook, and walkthrough for moving from a blank Specific Aims page to a submission-ready one-pager and budget justification in under three hours.
Drop
2026-W23
- ChatGPT
ChatGPT is the baseline general assistant card, but it now needs to be framed as a family of surfaces: consumer ChatGPT, Business/Enterprise workspaces, Codex/agent features, and the separate OpenAI API. Its strength is breadth — drafting, file work, multimodal analysis, coding, agents, and extensions — while its risk is that fluent output can make weak sourcing, privacy assumptions, or plan limits invisible.
Card - Claude
Claude is the careful-drafting and long-context card. The updated evidence reinforces its core lane: nuanced rewriting, document analysis, Projects, Artifacts, Research, Claude Code, Cowork, web search, and organization search/connectors. The editorial risk is quota and surface confusion: consumer Pro/Max, Team, Enterprise, and API/tool pricing are meaningfully different products.
Card - Connected Papers
Connected Papers remains the clean visual literature-map card. It is especially useful for orienting around one seed paper, but its core limitation is still methodological: a graph is not a reproducible search strategy. This pass adds official plan/feature detail, group/library plans, payment processing, scholarship support, and free-quota boundaries.
Card - Consensus
Consensus remains the claim-shaped academic Q&A card. It is strongest when the user has a research question and needs cited orientation quickly. This pass adds current corpus/user-scale claims, Research Agent, account/privacy controls, institutional full-text linking, and a sharper “answer box is not a literature review” warning.
Card - Crossref
Crossref is scholarly infrastructure, not a shiny discovery app. It remains the DOI metadata verification card: use it for DOI records, member/publisher metadata, references, funder/license fields, relation metadata, and citation plumbing. This pass adds the 2025/2026 REST API rate-limit shift and the practical “polite pool” guidance that makes Crossref usable in responsible pipelines.
Card - Cursor
Cursor remains the AI-native code editor flagship, but the card now needs to emphasize three separations: editor workflow versus headless agents, individual usage pools versus team/enterprise governance, and Privacy Mode versus ordinary provider transit. Cursor is excellent for repo-local Tab, Chat, Agent, MCP, cloud agents, and Bugbot workflows, but only if diffs are reviewed and tests run.
Card - DeepSeek
DeepSeek is the low-cost reasoning/API card. This pass sharpens the key facts: DeepSeek’s API is OpenAI- and Anthropic-compatible, current V4-Flash/V4-Pro pricing is extremely low, context length is listed at 1M, and legacy deepseek-chat / deepseek-reasoner names are being deprecated. The low price is real, but the editorial warning is bigger than usual: privacy, residency, institutional policy, and geopolitical availability need explicit review.
Card - ElevenLabs
ElevenLabs is the voice and audio AI card. This pass uses primary ElevenLabs pricing and privacy pages plus compliance/search results. The product’s quality can make outputs feel finished, but the editorial review must center voice consent, identity risk, disclosure, credit burn, and whether uploaded voices/interviews belong in a cloud voice system.
Card - Elicit
Elicit is no longer just a paper-search assistant; it is positioning itself as an evidence-synthesis system for reports, systematic reviews, dynamic screening, extraction, and sentence-level citations over a very large scientific corpus. The richer data supports a higher-confidence card, but the recommendation still depends on benchmarked extraction accuracy against a hand-coded table.
Card - Exa
Exa is one of the cleanest “search as infrastructure” cards in the VerdictPal stack. The current pricing and docs make the value clearer: real-time search, webpage text/highlights, configurable latency, contents retrieval, Answer endpoint, Deep Search, Monitors, and Agent runs. The core risk is economic and privacy-boundary creep when agents fan out across many searches, contents calls, summaries, and enrichment steps.
Card - Factory
Factory is the agent-native software-delivery card. It is strongest when the task is bigger than autocomplete: repo understanding, spec-planned refactors, terminal/IDE delegation, background agents, SDK use, and CI-style Droid Exec workflows. This pass adds current 2026 Factory/Droid source detail and sharpens the enterprise/security boundary.
Card - Firecrawl
Firecrawl is the web-extraction infrastructure card. It should be described as context extraction for agents and research pipelines, not as a generic search engine. This pass adds current credit economics, enterprise ZDR/SOC2 claims, endpoint credit costs, and the important warning that hosted scraping is always a trust, permission, and site-terms problem.
Card - Gamma
Gamma is the AI presentation and visual-document first-draft card. This pass uses official Gamma pricing/product/privacy/search snippets, but the pricing and privacy pages timed out on full load, so the dossier clearly marks what is verified from retrieved official snippets and what needs final live-page verification before publication.
Card - Gemini
Deepened from official Google pages: pricing across consumer (AI Plus/Pro/Ultra), Workspace add-on, Gemini API paid tier, and Vertex AI; privacy language for consumer vs API vs Workspace/Vertex; capability list refreshed to current Gemini app + API surface.
Card - Google Scholar
Google Scholar remains the academic-search baseline: broad, familiar, fast, and useful for finding known items, PDFs, library links, citation trails, alerts, and case law. Its weakness is methodological opacity. It should be recommended as a discovery and citation-chasing habit, not as a reproducible search strategy by itself.
Card - Grammarly
Grammarly is the revision-pass card, not the argument-writing card. This pass updates the card for Grammarly Pro and Superhuman-suite positioning: Free gives core correction plus 100 AI prompts, Pro adds rewrites/tone/brand features and 2,000 AI prompts, team features run up to 149 seats, and larger organizations move to Superhuman Go/enterprise-style controls. The core warning remains voice flattening and cloud text processing.
Card - Gumloop
Gumloop is the no-code AI workflow and agent-automation card. This pass is based on primary Gumloop pricing and privacy pages, plus official product/search results. The review should focus on whether credit-based agent workflows stay understandable after real data, credentials, retries, and handoffs enter the loop.
Card - Kagi
Kagi is the paid-search card: a user-funded, ad-free search engine with serious privacy engineering and enough Assistant functionality to be useful without turning into an AI slop layer. The strongest angle is incentive alignment — no ads, no result-click tracking, no analytics/telemetry on the main site, and paid plans that make the business model legible. The recommendation still depends on a 25-query search-quality benchmark against Google, Brave, DuckDuckGo, Perplexity, and You.com.
Card - Litmaps
Litmaps is the citation-map plus monitoring card. It is valuable when a user starts from seed papers and needs to discover connected work, visualize the field, sync references, and keep watch for new papers. This pass fixes the malformed sources area and adds current Litmaps feature, coverage, and DPA details.
Card - Mendeley
Mendeley is the Elsevier-backed cloud-sync reference manager with increasingly AI-shaped library features. The richer data clarifies the buyer’s question: Mendeley is capable for PDF import, annotation, Word citation, groups, and AI over a library, but the deciding tradeoff is Elsevier cloud trust, storage/AI plan economics, and future switching cost versus Zotero and Paperpile.
Card - Microsoft Copilot
Microsoft Copilot is a family-name card, not a single-product card. This pass uses Microsoft Learn pages for Microsoft 365 Copilot and Copilot Chat data protection plus Microsoft pricing search results. The key editorial job is to separate consumer Copilot, Microsoft 365 Copilot Chat, paid Microsoft 365 Copilot, and Azure/OpenAI developer paths.
Card - Mistral Le Chat
Mistral Le Chat is the EU-headquartered assistant comparison card. This pass strengthens the privacy and API boundary: Le Chat inputs/outputs are retained until account/conversation deletion, API inputs/outputs are generally kept for 30 rolling days for abuse monitoring unless zero data retention is activated, and Agents/Fine-tuning have separate retention rules. That makes Mistral useful to compare against US-default assistants, but not automatically compliant.
Card - n8n
n8n is the research-ops automation card. It is strongest when a team needs visible workflow glue across APIs, webhooks, databases, AI steps, Notion, and alerts. The updated evidence makes the trade-off clearer: every plan now leans into unlimited users/workflows/steps, but pricing is based on full workflow executions, and self-hosted still means the operator owns patching, backups, secrets, telemetry choices, and failure recovery.
Card - NotebookLM
NotebookLM remains the source-grounded workspace flagship. The strongest VerdictPal framing is not “AI search,” but “corpus-first reading workspace”: upload or discover sources, ask grounded questions, generate study artifacts, and verify claims against citations. This pass adds current Enterprise, Workspace, and Google AI plan distinctions so the privacy story is clearer.
Card - Obsidian
Obsidian is the local-first knowledge-base card. The important update is that Obsidian is now free for work as well as personal use, while paid services remain optional add-ons for Sync and Publish. That strengthens the card’s editorial angle: Obsidian is not a cloud workspace trying to lock your notes away; it is a Markdown vault where the user owns the files, but also owns the backup, plugin, and structure decisions.
Card - OpenAlex
OpenAlex is the open scholarly graph card. It sits between Crossref and Semantic Scholar: broader discovery and mapping than DOI registry metadata, more reproducible and open than closed search products, but still an aggregated graph that requires verification for final citations. This pass adds the 2026 usage-based pricing and privacy promise details.
Card - OpenRouter
OpenRouter is the model-router card. It is useful when a developer wants one OpenAI-compatible API surface for many LLMs, fallback routing, model comparisons, BYOK, and spend controls. The updated card sharpens the key caveat: routing convenience adds a platform fee and a subprocessor/provider-retention surface that must be controlled before confidential research or user data flows through it.
Card - Overleaf
Overleaf is the collaborative LaTeX card. This pass updates the card with current plan evidence, including AI allowance language, collaborator limits, 24x compile timeout on paid tiers, real-time track changes, and the subscriber-owned project entitlement model. The core trade remains: Overleaf removes local TeX and collaboration friction, but your project source and compile workflow live in a cloud editor.
Card - Paperpile
Paperpile is the convenience-first reference manager for Google Docs and browser-native writing. The stronger data now makes the tradeoff clearer: it has excellent Google Docs and PDF workflow coverage, annual pricing only, unlimited PDF storage on paid plans, and a cloud/Google-account privacy boundary that should be explicit for thesis and institutional work.
Card - Perplexity
Perplexity remains the cited-answer benchmark for consumer AI search, but the card now needs to treat Perplexity as two related products: the consumer answer engine and the API platform. The consumer product is useful for fast source maps and Deep Research drafts; the API platform is a separate developer surface with Sonar, Search, Agent, and Embeddings APIs. The hard editorial rule stays the same: citations are leads, not bibliography-ready evidence.
Card - PostHog
PostHog is the product-analytics and experimentation-stack card. It belongs in VerdictPal because product analytics, session replay, feature flags, experiments, surveys, pipelines, and AI/product ops determine how the product learns after launch. The recommendation stays gated because PostHog’s biggest risks are exactly the ones product teams underestimate: event volume, replay privacy, retention, and usage-based cost.
Card - Rayyan
Rayyan is a focused systematic-review screening workspace with stronger plan mechanics and AI boundaries than the earlier card captured. The product’s editorial value is title/abstract screening, deduplication, PICO extraction, PRISMA support, reviewer/viewer collaboration, mobile work, and institutional ResearchPilot. The gold-standard warning is unchanged: AI can assist review logistics, but the final inclusion/exclusion decision must remain human and auditable.
Card - Replit
Replit is the browser-based AI build-and-deploy card. It is useful when a user wants one place for ideation, agent building, database, hosting, collaboration, and publishing. The review must focus on code quality, credit burn, deployment security, and whether the user understands what the agent built.
Card - Research Rabbit
Research Rabbit is the visual exploration and “follow the trail” card. It belongs beside Connected Papers and Litmaps, but its emphasis is adaptive exploration, collections, and seeing how papers/authors/concepts connect. This pass adds current feature language, data-source hints, and a clearer privacy/method boundary.
Card - Scholarcy
Scholarcy is the paper-skimming and structured-summary card. It can save time for triage, accessibility, and organization, but it must be reviewed as a reading aid, not a reading replacement. This pass adds current product/pricing/help details, browser-extension behavior, export/library capabilities, and a sharper copyright/privacy boundary for uploaded PDFs.
Card - Scira AI
Scira remains the VerdictPal flagship template because it combines a cited search UI, open-source AGPL codebase, hosted Free/Pro/Max plans, an API platform, MCP surface, and a deep dossier structure. This pass replaces stale/malformed page content with a cleaner benchmarkable dossier while preserving the core stance: strong public recommendation, but benchmark-ready still false until the pending desk rows are actually run.
Card - Scite
Scite is the citation-context card. It is valuable because it asks a better question than raw citation count: did later papers support, contrast, or merely mention the claim? This pass strengthens the card with Scite’s current Smart Citations, MCP, API, publisher/full-text coverage, privacy-policy, and reference-check caveats. The public recommendation remains gated until label accuracy is benchmarked against actual citing passages.
Card - Semantic Scholar
Semantic Scholar remains a flagship academic-search layer because it combines a free scholarly search UI with TLDRs, author pages, alerts, a developer API, and downloadable scholarly graph data. It is more structured than Google Scholar and more paper-discovery oriented than Crossref. The warning is unchanged: AI summaries and graph edges are discovery aids, not evidence.
Card - Warp
Warp is the agentic terminal/workspace card. It is compelling because it puts local and cloud coding agents where developers already run commands, but the evaluation must center command safety, repo hygiene, and data controls.
Card - You.com
You.com should now be reviewed primarily as a web-search API platform for AI builders, not as a nostalgic consumer search/chat competitor. Its current public pitch centers on Web Search APIs, Research API, zero-data-retention options, SOC2, DPA readiness, high rate limits, and enterprise deployment. The card’s benchmark needs to test API source quality, freshness, cost, and data-control clarity against Exa, Perplexity Sonar/Search, Brave, Tavily-style APIs, and Kagi.
Card - Zotero
Zotero remains the strongest local-first default for student and academic reference management. The new data points reinforce the core positioning: the desktop workflow works without an account, sync is optional and disabled by default, paid storage funds a nonprofit project, and the main editorial risk is not capability but governance around sync, attachments, plugins, and group ownership.
Card - ChatPRD
Imported Notion state-of-art research and lowered the editorial fit to reflect weak desk usefulness plus Teams-gated Linear integration.
Card - Dark mode rebuilt for the evidence instrument
Dark surfaces now use muted amber, readable borders, and component-specific overrides instead of neon cards and inverted cream panels.
Drop - Trusted-source brief workflow shipped
The Guides pillar now has a real student research path, tool stack, playbook, and showcase for turning broad prompts into source-backed briefs.
Drop - Claude Mythos Preview
Released Claude Mythos research preview — tops Vals (73.42%) and SWE-bench Verified (93.9%, best-of-3).
Model
2026-W22
2026-W21
- Gemini 3.5 Flash
Released Gemini 3.5 Flash — MCP Atlas leader (83.6%), 1M context, cheapest frontier pricing.
Model
2026-W18
2026-W17
2026-W16
- Claude Opus 4.7
Released Claude Opus 4.7 — 13% coding lift, 3x production tasks, 3.75 MP vision.
Model
2026-W15
- GLM-5.1
Released GLM-5.1 with 200K context, 128K max output, long-horizon coding focus, and MIT-licensed open weights.
Model
2026-W14
2026-W12
2026-W11
2026-W10
2026-W08
2026-W07
2026-W06
- Claude Opus 4.6
Released Claude Opus 4.6 — first Opus with 1M context, 128K output, agent teams.
Model