What changed
2026-W36
- Claude Fable 5.1
Live Vals Index 68.83%. Live AA Intelligence Index v4.2 is 57 at max + default fallback.
Model - Gemini 3.7 Flash
Live AA Intelligence Index v4.2 is 45 at high effort. Gemini 3.8 Flash is now the current Flash; 3.7 stays listed for the faster 305 t/s endpoint.
Model - GPT-5.6 Sol
Live AA Intelligence Index v4.2 is 51 at max. GPT-6 Astra is the current OpenAI flagship; Sol stays the cheaper GPT-5.6 line.
Model - Grok 4.6
Live AA Intelligence Index v4.2 is 51 at high effort; AA now lists a 500k context window.
Model - September 2026 frontier: GPT-6 Astra, Gemini 3.8 Flash, Muse Spark 1.3, Hy4 preview
OpenAI's 3 Sep GPT-6 Astra, Google's 2 Sep Gemini 3.8 Flash, and Meta's 2 Sep Muse Spark 1.3 are now on the atlas with live Artificial Analysis Intelligence Index v4.2 scores. Tencent's Hy4 preview is recorded from the official launch. Claude Fable 5.1 still leads the live Vals Index at 68.83%. /models
Drop - GPT-6 Astra
GPT-6 Astra GA on the OpenAI API and a phased ChatGPT / AWS rollout. List $10 / $50 per 1M tokens.
Model - Gemini 3.8 Flash
Gemini 3.8 Flash generally available. AA high-effort Intelligence Index v4.2 is 47; intro price unchanged from 3.7 Flash.
Model - Muse Spark 1.2
Muse Spark 1.3 is now the current Muse Spark line. 1.2 stays listed for the v4.1 xhigh snapshot.
Model - Muse Spark 1.3
Muse Spark 1.3 available in Muse Code and the Meta Model API. AA max-effort Intelligence Index v4.2 is 53.
Model - AA-Omniscience and CritPt join the benchmark atlas
Two missing slices of AA Intelligence Index v4.1 now have explainers and sourced rows: AA-Omniscience (factual recall vs hallucination — the knowledge primary) and CritPt (research-level physics, still under 40% at the frontier). /benchmarks/aa-omniscience · /benchmarks/critpt
Drop - Claude Fable 5.1
Claude Fable 5.1 GA on the Claude API, Bedrock, Vertex, and Microsoft Foundry. Cache reads cut to $0.25 / 1M tokens.
Model - Claude Fable 5.1 is on the model atlas — Vals Index 67.87%
Anthropic's 1 Sep Mythos-class refresh of Fable 5. Same $10 / $50 list, cache reads cut to $0.25 per 1M tokens. Vals puts it at 67.87% on the Index, ahead of Opus 5 and Fable 5. AA Intelligence Index is not in the 1 Sep API snapshot yet. Compare it with GPT-5.6 Sol or Fable 5. /models/claude-fable-5-1
Drop
2026-W35
- Hy4 preview
Hy4 preview open-sourced and listed on TokenHub / OpenRouter at $0.834 / $2.501 per 1M tokens.
Model - Terminal-Bench-Science 0.1 is on the atlas — best row 30%
Stanford and Laude's 70-task science-agent suite: real lab workflows, still far from ceiling. Claude Opus 5 + Claude Code leads at 30.0%; GPT-5.6 Sol + Codex is 22.4%; GLM-5.3 is the strongest open-weight row at 8.1%. Not Terminal-Bench 2.1. /benchmarks/terminal-bench-science
Drop - Three academic guide trios, rebuilt
Student researcher, literature reviewer, and peer reviewer — each with a matching stack and playbook. Peer review starts with the venue AI policy. Retired grant, coding, and deck paths redirect to /guides.
Drop - GLM-5.3-Flash is on the model atlas — AA index 57.5
Z.ai's Flash sibling of GLM-5.3, from the 1 Sep Artificial Analysis snapshot: Intelligence Index 57.5 at $0.15 / $0.50 per 1M tokens. Two points behind GLM-5.3 max at a tenth of the list price. /models/glm-5-3-flash · /compare/models/glm-5-3-vs-glm-5-3-flash
Drop - GLM-5.3 is on the model atlas — AA index 59.5
Z.ai's August coding refresh is on the atlas from the 2026-08-24 Artificial Analysis snapshot: max-effort Intelligence Index 59.5, API list $1.40 / $4.40 per 1M tokens. Same base as GLM-5.2; weights were still unpublished on this check. Ledger rows sync from AA; editorial depth still needs a desk pass. Compare it with Qwen3.8 27B. /models/glm-5-3 · /compare/models/glm-5-3-vs-qwen3-8-27b
Drop - Qwen3.8 27B is on the model atlas — AA index 52
Alibaba's dense open-weight Qwen3.8 sibling is on the atlas from the 2026-08-24 Artificial Analysis snapshot: xhigh Intelligence Index 52, hosted list $0.50 / $3 per 1M tokens, Apache 2.0 on Hugging Face. This is not hosted Qwen3.8 Max. Ledger rows sync from AA; editorial depth still needs a desk pass. Compare it with Qwen3.8 Max. /models/qwen3-8-27b · /compare/models/qwen3-8-27b-vs-qwen3-8-max
Drop
2026-W34
- GLM-5.3
Hosted API listed on Artificial Analysis (max-effort index 59.5, $1.40 / $4.40 per 1M tokens).
Model
2026-W33
- Qwen3.8 2.4T is on the model atlas — AA index 57.7
Alibaba's open-weight Qwen-Max-class MoE (2.4T total / 95B active) is on the atlas from the 2026-08-15 Artificial Analysis snapshot: Intelligence Index 57.7, hosted list $2 / $6 per 1M tokens. This is the published checkpoint, not hosted Qwen3.8 Max. Ledger rows sync from AA; editorial depth still needs a desk pass. Compare it with Qwen3.8 Max. /models/qwen3-8-2-4t-a95b · /compare/models/qwen3-8-2-4t-a95b-vs-qwen3-8-max
Drop - Qwen3.8 27B
Open weights published. AA Intelligence Index 52 on xhigh; hosted list $0.50 / $3 per 1M tokens.
Model - Gemini 3.6 Flash (high)
Objective refresh: DeepMind card 1M context / 64K output; Google intro price $0.75 / $3.75; AA high-effort index 51.6.
Model - Gemini 3.7 Flash
Gemini 3.7 Flash generally available. AA high-effort Intelligence Index 56; intro price $0.75 / $3.75 per 1M tokens. Google reports gains over 3.6 Flash on FrontierCode 1.1, DeepSWE v1.1, WebDev Arena, and GDP.pdf — vendor-claimed, not ledger rows.
Model - Gemini 3.7 Flash is on the model atlas — AA index 56
Google's August Flash workhorse is on the atlas from the Artificial Analysis snapshot: high-effort Intelligence Index 56, intro price $0.75 / $3.75 per 1M tokens through 31 Dec 2026. Ledger rows sync from AA; editorial depth still needs a desk pass. Compare it with 3.6 Flash or Claude Sonnet 4.6. /models/gemini-3-7-flash · /compare/models/gemini-3-6-flash-vs-gemini-3-7-flash
Drop - Grok 4.6
Grok 4.6 appears on Artificial Analysis (high-effort index 60.9, $2 / $6 per 1M tokens).
Model - Grok 4.6 is on the model atlas — AA index 60.9
xAI's August flagship is on the atlas from the Artificial Analysis snapshot: high-effort Intelligence Index 60.9, $2 / $6 per 1M tokens. Ledger rows sync from AA; editorial depth still needs a desk pass. Compare it with GPT-5.6 Sol. /models/grok-4-6 · /compare/models/gpt-5-6-sol-vs-grok-4-6
Drop - Qwen3.8 2.4T A95B
Open weights published. AA Intelligence Index 57.7; hosted list $2 / $6 per 1M tokens.
Model - Nemotron 3.5 Lightning
Nemotron 3.5 Lightning on Artificial Analysis (index 23.6, $0.05 / $0.20).
Model - Muse Glimmer
Muse Glimmer open-weighted (Apache 2.0, 30B). AA high-effort index 35.1; hosted list $0.325 / $1.35 per 1M tokens.
Model
2026-W32
- Muse Spark 1.2
Muse Spark 1.2 appears on Artificial Analysis (xhigh index 56.8, $1.25 / $4.25 per 1M tokens).
Model - VerdictPal is live, and the Atlas Report puts numbers on the whole atlas
The launch note is published, plus a new original-research surface: the Atlas Report computes aggregate pricing, privacy, benchmark-saturation and freshness figures from every published tool, and it is free to cite with a link. Highlights this edition — most dossiers do not record any stance on training with your content, paid entry plans cluster around the median, and frontier model input prices spread by more than three orders of magnitude. /report · /launch
Drop
2026-W29
- GPT-5.6 Luna
Recorded general availability across ChatGPT, Codex, and the API and refreshed benchmark evidence.
Model - GPT-5.6 Sol
Recorded general availability across ChatGPT, Codex, and the API, then refreshed Vals and Artificial Analysis benchmark evidence.
Model - GPT-5.6 Terra
Recorded general availability across ChatGPT, Codex, and the API and refreshed benchmark evidence.
Model
2026-W28
- AI Builder Design Test on the Lab bench — four builders, same prompts, very different results
VerdictPal Lab's second first-party protocol: four identical multi-turn design prompts through Lovable, v0, Replit, and Bolt. Run 01 recorded a Turn 1 ranking (Lovable → v0 → Replit → Bolt) and the cost-of-arrival in each vendor's native unit — no padded leaderboard, formal scored rows pending a repeatable rubric. /benchmarks/ai-builder-design-test · /benchmarks/lab
Drop - Citation Fidelity v0.1 protocol frozen on the Lab bench
VerdictPal Lab's first first-party benchmark: twelve research questions, dual-reviewer citation grading, frozen runbook on the atlas page — protocol published, desk run still pending, no padded leaderboard. /benchmarks/citation-fidelity · /benchmarks/lab
Drop - Claude
Evidence integrity fix: reverted false editorial-hands-on upgrade; evidence stays synthesized and qualityGate downgraded flagship → solid per the editorial evidence ladder. No desk hands-on trial on file; benchmarkReady remains false.
Tool
2026-W27
- Five research guides now on the atlas
Citation audit, systematic review protocol, data cleaning pipeline, grant writer path, and peer reviewer path — solid guides with step blocks and verification checklists.
Drop - GPT-5.6 Sol
Updated public pricing: Sol is listed at $5 input / $30 output per 1M tokens, with 1.25x cache-write billing and 90% cached-read discount.
Model - GPT-5.6 Terra
Updated public pricing: Terra is listed at $2.50 input / $15 output per 1M tokens.
Model - OpenRouter
Added LongCat-2.0/Owl Alpha release context: Meituan identified the OpenRouter stealth route as LongCat-backed after it had already become a high-volume coding model. Treat it as evidence of routing-market adoption, not as a privacy shortcut or independently verified benchmark row.
Tool - LongCat-2.0
Meituan revealed LongCat-2.0 and identified OpenRouter's Owl Alpha stealth route as LongCat-backed; GitHub/Hugging Face pages say model weights are coming soon.
Model
2026-W26
- Student research workflow on the guides pillar
Path, stack, playbook, and showcase for turning broad prompts into source-backed briefs — the workflow we point new readers at first.
Drop - Factory flagship dossier shipped
Agent-native delivery with Droid, Spec Mode, missions, and enterprise controls — fit 74 with the autonomy and pricing caveats on the record.
Drop - Cursor flagship dossier shipped
VS Code-compatible AI editor at fit 81 — Tab, Chat, Agent, Privacy Mode, team governance, and the diff-review failure modes we publish before you merge.
Drop - Gemini flagship dossier shipped
Google's multimodal assistant ecosystem at fit 75 — Search grounding, Deep Research, storage bundles, and where it is not a cited search engine.
Drop
2026-W25
- NotebookLM flagship dossier shipped
Source-grounded reading workspace at fit 78 — corpus chat, Audio Overviews, Flashcards, and the empty-corpus trap documented on the dossier.
Drop - Perplexity flagship dossier shipped
Cited search front door at fit 72 — consumer Pro/Max, Deep Research, Sonar APIs, and failure modes for bibliographies you have not opened.
Drop - Cursor
Production-readiness audit pass: editorsNote and pricing/privacy checked dates confirmed. benchmarkReady remains false intentionally until agent-diff safety, usage-burn, and Privacy Mode runs land.
Tool - Factory
Production-readiness audit pass: editorsNote and pricing/privacy checked dates confirmed. benchmarkReady remains false intentionally until Spec Mode, Droid Exec autonomy, and enterprise-control runs land.
Tool - Gemini
Production-readiness audit pass: editorsNote and pricing/privacy checked dates confirmed. No editorial claims changed; bench rows remain explicitly pending until the desk protocol runs (Next action in qualityGate.notes).
Tool - Perplexity
Production-readiness audit pass: editorsNote, pricing/privacy checked dates confirmed; Sonar/Search API token + per-request fee rows aligned with docs.perplexity.ai. benchmarkReady remains false intentionally until desk citation audit and Sonar latency runs land.
Tool - Command Code
Shipped Command Code flagship tool — terminal agent with taste learning, open-model deals, and full pricing/privacy research from commandcode.ai.
Tool - Command Code flagship dossier shipped
Terminal coding agent with taste-1 learning, open-model deals from $1/mo, and the full pricing and privacy picture — including what lives in .commandcode/.
Drop - OpenCode
Shipped OpenCode flagship tool — OSS agent, Go/Zen pricing, Plan/Build modes, and Zen privacy exceptions documented.
Tool - OpenCode flagship dossier shipped
The open-source coding agent at fit 95 — Plan/Build modes, Go and Zen pricing, 75+ providers, and the free-model privacy caveats on the record.
Drop - Raycast
Shipped Raycast as a flagship Knowledge work tool — the macOS launcher we would install on every research Mac before debating another AI tab.
Tool - Raycast flagship dossier shipped
The macOS launcher we install before debating another AI tab — extensions, Quick AI, clipboard, and honest failure modes at fit 97.
Drop
2026-W24
- GLM-5.2
Released GLM-5.2 with 1M context, 128K max output, two thinking modes, and long-horizon coding focus. MIT open weights announced for the following week.
Model - Kimi K2.7 Code
Released Kimi K2.7 Code with 256K context, mandatory thinking mode, multimodal tool examples, and reported 30% lower reasoning-token usage versus K2.6.
Model - Canva Business
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - ChatGPT
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - ChatPRD
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Connected Papers
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Consensus
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Crossref
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - DeepSeek
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - ElevenLabs
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Elicit
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Exa
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Firecrawl
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Gamma
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Google Scholar
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Grammarly
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Gumloop
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Kagi
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Linear
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Litmaps
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Lovable
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Magic Patterns
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Manus
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Mendeley
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Microsoft Copilot
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Mistral Le Chat
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Mobbin
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - n8n
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Obsidian
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - OpenAlex
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - OpenRouter
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Overleaf
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Paperpile
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - PostHog
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Rayyan
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Replit
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Research Rabbit
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Scholarcy
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Scira AI
Refreshed pricing evidence from scira.ai/about and api.scira.ai/pricing, added machine-readable pricing tiers, and marked benchmark evidence ready for the public dossier while keeping individual desk benchmark rows explicit about what still needs live runs.
Tool - Scite
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Semantic Scholar
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Warp
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Wispr Flow
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - You.com
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Zotero
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Tool - Grant deadline set (archived)
The grant-deadline path has been retired. Academic guides now live as three trios: student researcher, literature reviewer, and peer reviewer.
Drop
2026-W23
- ChatGPT
ChatGPT is the baseline general assistant tool, but it now needs to be framed as a family of surfaces: consumer ChatGPT, Business/Enterprise workspaces, Codex/agent features, and the separate OpenAI API. Its strength is breadth — drafting, file work, multimodal analysis, coding, agents, and extensions — while its risk is that fluent output can make weak sourcing, privacy assumptions, or plan limits invisible.
Tool - Claude
Claude is the careful-drafting and long-context tool. The updated evidence reinforces its core lane: nuanced rewriting, document analysis, Projects, Artifacts, Research, Claude Code, Cowork, web search, and organization search/connectors. The editorial risk is quota and surface confusion: consumer Pro/Max, Team, Enterprise, and API/tool pricing are meaningfully different products.
Tool - Connected Papers
Connected Papers remains the clean visual literature-map card. It is especially useful for orienting around one seed paper, but its core limitation is still methodological: a graph is not a reproducible search strategy. This pass adds official plan/feature detail, group/library plans, payment processing, scholarship support, and free-quota boundaries.
Tool - Consensus
Consensus remains the claim-shaped academic Q&A tool. It is strongest when the user has a research question and needs cited orientation quickly. This pass adds current corpus/user-scale claims, Research Agent, account/privacy controls, institutional full-text linking, and a sharper “answer box is not a literature review” warning.
Tool - Crossref
Crossref is scholarly infrastructure, not a shiny discovery app. It remains the DOI metadata verification card: use it for DOI records, member/publisher metadata, references, funder/license fields, relation metadata, and citation plumbing. This pass adds the 2025/2026 REST API rate-limit shift and the practical “polite pool” guidance that makes Crossref usable in responsible pipelines.
Tool - Cursor
Cursor remains the AI-native code editor flagship, but the tool now needs to emphasize three separations: editor workflow versus headless agents, individual usage pools versus team/enterprise governance, and Privacy Mode versus ordinary provider transit. Cursor is excellent for repo-local Tab, Chat, Agent, MCP, cloud agents, and Bugbot workflows, but only if diffs are reviewed and tests run.
Tool - DeepSeek
DeepSeek is the low-cost reasoning/API tool. This pass sharpens the key facts: DeepSeek’s API is OpenAI- and Anthropic-compatible, current V4-Flash/V4-Pro pricing is extremely low, context length is listed at 1M, and legacy deepseek-chat / deepseek-reasoner names are being deprecated. The low price is real, but the editorial warning is bigger than usual: privacy, residency, institutional policy, and geopolitical availability need explicit review.
Tool - ElevenLabs
ElevenLabs is the voice and audio AI tool. This pass uses primary ElevenLabs pricing and privacy pages plus compliance/search results. The product’s quality can make outputs feel finished, but the editorial review must center voice consent, identity risk, disclosure, credit burn, and whether uploaded voices/interviews belong in a cloud voice system.
Tool - Elicit
Elicit is no longer just a paper-search assistant; it is positioning itself as an evidence-synthesis system for reports, systematic reviews, dynamic screening, extraction, and sentence-level citations over a very large scientific corpus. The richer data supports a higher-confidence card, but the recommendation still depends on benchmarked extraction accuracy against a hand-coded table.
Tool - Exa
Exa is one of the cleanest “search as infrastructure” cards in the VerdictPal stack. The current pricing and docs make the value clearer: real-time search, webpage text/highlights, configurable latency, contents retrieval, Answer endpoint, Deep Search, Monitors, and Agent runs. The core risk is economic and privacy-boundary creep when agents fan out across many searches, contents calls, summaries, and enrichment steps.
Tool - Factory
Factory is the agent-native software-delivery tool. It is strongest when the task is bigger than autocomplete: repo understanding, spec-planned refactors, terminal/IDE delegation, background agents, SDK use, and CI-style Droid Exec workflows. This pass adds current 2026 Factory/Droid source detail and sharpens the enterprise/security boundary.
Tool - Firecrawl
Firecrawl is the web-extraction infrastructure tool. It should be described as context extraction for agents and research pipelines, not as a generic search engine. This pass adds current credit economics, enterprise ZDR/SOC2 claims, endpoint credit costs, and the important warning that hosted scraping is always a trust, permission, and site-terms problem.
Tool - Gamma
Gamma is the AI presentation and visual-document first-draft tool. This pass uses official Gamma pricing/product/privacy/search snippets, but the pricing and privacy pages timed out on full load, so the dossier clearly marks what is verified from retrieved official snippets and what needs final live-page verification before publication.
Tool - Gemini
Deepened from official Google pages: pricing across consumer (AI Plus/Pro/Ultra), Workspace add-on, Gemini API paid tier, and Vertex AI; privacy language for consumer vs API vs Workspace/Vertex; capability list refreshed to current Gemini app + API surface.
Tool - Google Scholar
Google Scholar remains the academic-search baseline: broad, familiar, fast, and useful for finding known items, PDFs, library links, citation trails, alerts, and case law. Its weakness is methodological opacity. It should be recommended as a discovery and citation-chasing habit, not as a reproducible search strategy by itself.
Tool - Grammarly
Grammarly is the revision-pass tool, not the argument-writing tool. This pass updates the tool for Grammarly Pro and Superhuman-suite positioning: Free gives core correction plus 100 AI prompts, Pro adds rewrites/tone/brand features and 2,000 AI prompts, team features run up to 149 seats, and larger organizations move to Superhuman Go/enterprise-style controls. The core warning remains voice flattening and cloud text processing.
Tool - Gumloop
Gumloop is the no-code AI workflow and agent-automation tool. This pass is based on primary Gumloop pricing and privacy pages, plus official product/search results. The review should focus on whether credit-based agent workflows stay understandable after real data, credentials, retries, and handoffs enter the loop.
Tool - Kagi
Kagi is the paid-search tool: a user-funded, ad-free search engine with serious privacy engineering and enough Assistant functionality to be useful without turning into an AI slop layer. The strongest angle is incentive alignment — no ads, no result-click tracking, no analytics/telemetry on the main site, and paid plans that make the business model legible. The recommendation still depends on a 25-query search-quality benchmark against Google, Brave, DuckDuckGo, Perplexity, and You.com.
Tool - Litmaps
Litmaps is the citation-map plus monitoring tool. It is valuable when a user starts from seed papers and needs to discover connected work, visualize the field, sync references, and keep watch for new papers. This pass fixes the malformed sources area and adds current Litmaps feature, coverage, and DPA details.
Tool - Mendeley
Mendeley is the Elsevier-backed cloud-sync reference manager with increasingly AI-shaped library features. The richer data clarifies the buyer’s question: Mendeley is capable for PDF import, annotation, Word citation, groups, and AI over a library, but the deciding tradeoff is Elsevier cloud trust, storage/AI plan economics, and future switching cost versus Zotero and Paperpile.
Tool - Microsoft Copilot
Microsoft Copilot is a family-name card, not a single-product card. This pass uses Microsoft Learn pages for Microsoft 365 Copilot and Copilot Chat data protection plus Microsoft pricing search results. The key editorial job is to separate consumer Copilot, Microsoft 365 Copilot Chat, paid Microsoft 365 Copilot, and Azure/OpenAI developer paths.
Tool - Mistral Le Chat
Mistral Le Chat is the EU-headquartered assistant comparison tool. This pass strengthens the privacy and API boundary: Le Chat inputs/outputs are retained until account/conversation deletion, API inputs/outputs are generally kept for 30 rolling days for abuse monitoring unless zero data retention is activated, and Agents/Fine-tuning have separate retention rules. That makes Mistral useful to compare against US-default assistants, but not automatically compliant.
Tool - n8n
n8n is the research-ops automation tool. It is strongest when a team needs visible workflow glue across APIs, webhooks, databases, AI steps, Notion, and alerts. The updated evidence makes the trade-off clearer: every plan now leans into unlimited users/workflows/steps, but pricing is based on full workflow executions, and self-hosted still means the operator owns patching, backups, secrets, telemetry choices, and failure recovery.
Tool - NotebookLM
NotebookLM remains the source-grounded workspace flagship. The strongest VerdictPal framing is not “AI search,” but “corpus-first reading workspace”: upload or discover sources, ask grounded questions, generate study artifacts, and verify claims against citations. This pass adds current Enterprise, Workspace, and Google AI plan distinctions so the privacy story is clearer.
Tool - Obsidian
Obsidian is the local-first knowledge-base tool. The important update is that Obsidian is now free for work as well as personal use, while paid services remain optional add-ons for Sync and Publish. That strengthens the tool’s editorial angle: Obsidian is not a cloud workspace trying to lock your notes away; it is a Markdown vault where the user owns the files, but also owns the backup, plugin, and structure decisions.
Tool - OpenAlex
OpenAlex is the open scholarly graph tool. It sits between Crossref and Semantic Scholar: broader discovery and mapping than DOI registry metadata, more reproducible and open than closed search products, but still an aggregated graph that requires verification for final citations. This pass adds the 2026 usage-based pricing and privacy promise details.
Tool - OpenRouter
OpenRouter is the model-router tool. It is useful when a developer wants one OpenAI-compatible API surface for many LLMs, fallback routing, model comparisons, BYOK, and spend controls. The updated tool sharpens the key caveat: routing convenience adds a platform fee and a subprocessor/provider-retention surface that must be controlled before confidential research or user data flows through it.
Tool - Overleaf
Overleaf is the collaborative LaTeX tool. This pass updates the tool with current plan evidence, including AI allowance language, collaborator limits, 24x compile timeout on paid tiers, real-time track changes, and the subscriber-owned project entitlement model. The core trade remains: Overleaf removes local TeX and collaboration friction, but your project source and compile workflow live in a cloud editor.
Tool - Paperpile
Paperpile is the convenience-first reference manager for Google Docs and browser-native writing. The stronger data now makes the tradeoff clearer: it has excellent Google Docs and PDF workflow coverage, annual pricing only, unlimited PDF storage on paid plans, and a cloud/Google-account privacy boundary that should be explicit for thesis and institutional work.
Tool - Perplexity
Perplexity remains the cited-answer benchmark for consumer AI search, but the tool now needs to treat Perplexity as two related products: the consumer answer engine and the API platform. The consumer product is useful for fast source maps and Deep Research drafts; the API platform is a separate developer surface with Sonar, Search, Agent, and Embeddings APIs. The hard editorial rule stays the same: citations are leads, not bibliography-ready evidence.
Tool - PostHog
PostHog is the product-analytics and experimentation-stack tool. It belongs in VerdictPal because product analytics, session replay, feature flags, experiments, surveys, pipelines, and AI/product ops determine how the product learns after launch. The recommendation stays gated because PostHog’s biggest risks are exactly the ones product teams underestimate: event volume, replay privacy, retention, and usage-based cost.
Tool - Rayyan
Rayyan is a focused systematic-review screening workspace with stronger plan mechanics and AI boundaries than the earlier tool captured. The product’s editorial value is title/abstract screening, deduplication, PICO extraction, PRISMA support, reviewer/viewer collaboration, mobile work, and institutional ResearchPilot. The gold-standard warning is unchanged: AI can assist review logistics, but the final inclusion/exclusion decision must remain human and auditable.
Tool - Replit
Replit is the browser-based AI build-and-deploy tool. It is useful when a user wants one place for ideation, agent building, database, hosting, collaboration, and publishing. The review must focus on code quality, credit burn, deployment security, and whether the user understands what the agent built.
Tool - Research Rabbit
Research Rabbit is the visual exploration and “follow the trail” tool. It belongs beside Connected Papers and Litmaps, but its emphasis is adaptive exploration, collections, and seeing how papers/authors/concepts connect. This pass adds current feature language, data-source hints, and a clearer privacy/method boundary.
Tool - Scholarcy
Scholarcy is the paper-skimming and structured-summary tool. It can save time for triage, accessibility, and organization, but it must be reviewed as a reading aid, not a reading replacement. This pass adds current product/pricing/help details, browser-extension behavior, export/library capabilities, and a sharper copyright/privacy boundary for uploaded PDFs.
Tool - Scira AI
Scira remains the VerdictPal flagship template because it combines a cited search UI, open-source AGPL codebase, hosted Free/Pro/Max plans, an API platform, MCP surface, and a deep dossier structure. This pass replaces stale/malformed page content with a cleaner benchmarkable dossier while preserving the core stance: strong public recommendation, but benchmark-ready still false until the pending desk rows are actually run.
Tool - Scite
Scite is the citation-context tool. It is valuable because it asks a better question than raw citation count: did later papers support, contrast, or merely mention the claim? This pass strengthens the tool with Scite’s current Smart Citations, MCP, API, publisher/full-text coverage, privacy-policy, and reference-check caveats. The public recommendation remains gated until label accuracy is benchmarked against actual citing passages.
Tool - Semantic Scholar
Semantic Scholar remains a flagship academic-search layer because it combines a free scholarly search UI with TLDRs, author pages, alerts, a developer API, and downloadable scholarly graph data. It is more structured than Google Scholar and more paper-discovery oriented than Crossref. The warning is unchanged: AI summaries and graph edges are discovery aids, not evidence.
Tool - Warp
Warp is the agentic terminal/workspace tool. It is compelling because it puts local and cloud coding agents where developers already run commands, but the evaluation must center command safety, repo hygiene, and data controls.
Tool - You.com
You.com should now be reviewed primarily as a web-search API platform for AI builders, not as a nostalgic consumer search/chat competitor. Its current public pitch centers on Web Search APIs, Research API, zero-data-retention options, SOC2, DPA readiness, high rate limits, and enterprise deployment. The tool’s benchmark needs to test API source quality, freshness, cost, and data-control clarity against Exa, Perplexity Sonar/Search, Brave, Tavily-style APIs, and Kagi.
Tool - Zotero
Zotero remains the strongest local-first default for student and academic reference management. The new data points reinforce the core positioning: the desktop workflow works without an account, sync is optional and disabled by default, paid storage funds a nonprofit project, and the main editorial risk is not capability but governance around sync, attachments, plugins, and group ownership.
Tool - Canva Business
Imported from the Notion Tool pipeline (Card pipeline) with Scira-gate status corrected to solid.
Tool - ChatPRD
Imported Notion state-of-art research and lowered the editorial fit to reflect weak desk usefulness plus Teams-gated Linear integration.
Tool - Linear
Imported from the Notion Tool pipeline (Card pipeline) with Scira-gate status corrected to solid.
Tool - Lovable
Imported from the Notion Tool pipeline (Card pipeline) with Scira-gate status corrected to solid.
Tool - Magic Patterns
Imported from the Notion Tool pipeline (Card pipeline) with Scira-gate status corrected to solid.
Tool - Mobbin
Imported from the Notion Tool pipeline (Card pipeline) with Scira-gate status corrected to solid.
Tool - Wispr Flow
Imported from the Notion Tool pipeline (Card pipeline) with Scira-gate status corrected to solid.
Tool - Dark mode rebuilt for the evidence instrument
Dark surfaces now use muted amber, readable borders, and component-specific overrides instead of neon panels and inverted cream panels.
Drop - Trusted-source brief workflow shipped
The Guides pillar now has a real student research path, tool stack, playbook, and showcase for turning broad prompts into source-backed briefs.
Drop - Claude Mythos Preview
Released Claude Mythos research preview — tops Vals (73.42%) and SWE-bench Verified (93.9%, best-of-3).
Model
2026-W22
2026-W21
- Gemini 3.5 Flash
Released Gemini 3.5 Flash — MCP Atlas leader (83.6%), 1M context, cheapest frontier pricing.
Model
2026-W18
2026-W17
2026-W16
- Claude Opus 4.7
Released Claude Opus 4.7 — 13% coding lift, 3x production tasks, 3.75 MP vision.
Model
2026-W15
- GLM-5.1
Released GLM-5.1 with 200K context, 128K max output, long-horizon coding focus, and MIT-licensed open weights.
Model
2026-W14
2026-W12
2026-W11
2026-W10
2026-W08
2026-W07
2026-W06
- Claude Opus 4.6
Released Claude Opus 4.6 — first Opus with 1M context, 128K output, agent teams.
Model