VerdictPal · editorial desk · updated 5 Sep 2026VerdictPal
Freshness

What changed

Tool changelog entries and Pack drops, grouped by week.

2026-W36

  1. Claude Fable 5.1

    Live Vals Index 68.83%. Live AA Intelligence Index v4.2 is 57 at max + default fallback.

    Model
  2. Gemini 3.7 Flash

    Live AA Intelligence Index v4.2 is 45 at high effort. Gemini 3.8 Flash is now the current Flash; 3.7 stays listed for the faster 305 t/s endpoint.

    Model
  3. GPT-5.6 Sol

    Live AA Intelligence Index v4.2 is 51 at max. GPT-6 Astra is the current OpenAI flagship; Sol stays the cheaper GPT-5.6 line.

    Model
  4. Grok 4.6

    Live AA Intelligence Index v4.2 is 51 at high effort; AA now lists a 500k context window.

    Model
  5. September 2026 frontier: GPT-6 Astra, Gemini 3.8 Flash, Muse Spark 1.3, Hy4 preview

    OpenAI's 3 Sep GPT-6 Astra, Google's 2 Sep Gemini 3.8 Flash, and Meta's 2 Sep Muse Spark 1.3 are now on the atlas with live Artificial Analysis Intelligence Index v4.2 scores. Tencent's Hy4 preview is recorded from the official launch. Claude Fable 5.1 still leads the live Vals Index at 68.83%. /models

    Drop
  6. GPT-6 Astra

    GPT-6 Astra GA on the OpenAI API and a phased ChatGPT / AWS rollout. List $10 / $50 per 1M tokens.

    Model
  7. Gemini 3.8 Flash

    Gemini 3.8 Flash generally available. AA high-effort Intelligence Index v4.2 is 47; intro price unchanged from 3.7 Flash.

    Model
  8. Gemini 3.8 Flash Cyber

    Gemini 3.8 Flash Cyber announced for Fairwind trusted defenders.

    Model
  9. Muse Spark 1.2

    Muse Spark 1.3 is now the current Muse Spark line. 1.2 stays listed for the v4.1 xhigh snapshot.

    Model
  10. Muse Spark 1.3

    Muse Spark 1.3 available in Muse Code and the Meta Model API. AA max-effort Intelligence Index v4.2 is 53.

    Model
  11. AA-Omniscience and CritPt join the benchmark atlas

    Two missing slices of AA Intelligence Index v4.1 now have explainers and sourced rows: AA-Omniscience (factual recall vs hallucination — the knowledge primary) and CritPt (research-level physics, still under 40% at the frontier). /benchmarks/aa-omniscience · /benchmarks/critpt

    Drop
  12. Claude Fable 5

    Vals Index 66.04% on the 1 Sep snapshot; Fable 5.1 takes the lead at 67.87%.

    Model
  13. Claude Fable 5.1

    Claude Fable 5.1 GA on the Claude API, Bedrock, Vertex, and Microsoft Foundry. Cache reads cut to $0.25 / 1M tokens.

    Model
  14. Claude Fable 5.1 is on the model atlas — Vals Index 67.87%

    Anthropic's 1 Sep Mythos-class refresh of Fable 5. Same $10 / $50 list, cache reads cut to $0.25 per 1M tokens. Vals puts it at 67.87% on the Index, ahead of Opus 5 and Fable 5. AA Intelligence Index is not in the 1 Sep API snapshot yet. Compare it with GPT-5.6 Sol or Fable 5. /models/claude-fable-5-1

    Drop

2026-W35

  1. Hy4 preview

    Hy4 preview open-sourced and listed on TokenHub / OpenRouter at $0.834 / $2.501 per 1M tokens.

    Model
  2. Terminal-Bench-Science 0.1 is on the atlas — best row 30%

    Stanford and Laude's 70-task science-agent suite: real lab workflows, still far from ceiling. Claude Opus 5 + Claude Code leads at 30.0%; GPT-5.6 Sol + Codex is 22.4%; GLM-5.3 is the strongest open-weight row at 8.1%. Not Terminal-Bench 2.1. /benchmarks/terminal-bench-science

    Drop
  3. Three academic guide trios, rebuilt

    Student researcher, literature reviewer, and peer reviewer — each with a matching stack and playbook. Peer review starts with the venue AI policy. Retired grant, coding, and deck paths redirect to /guides.

    Drop
  4. GLM-5.3-Flash is on the model atlas — AA index 57.5

    Z.ai's Flash sibling of GLM-5.3, from the 1 Sep Artificial Analysis snapshot: Intelligence Index 57.5 at $0.15 / $0.50 per 1M tokens. Two points behind GLM-5.3 max at a tenth of the list price. /models/glm-5-3-flash · /compare/models/glm-5-3-vs-glm-5-3-flash

    Drop
  5. GLM-5.3 is on the model atlas — AA index 59.5

    Z.ai's August coding refresh is on the atlas from the 2026-08-24 Artificial Analysis snapshot: max-effort Intelligence Index 59.5, API list $1.40 / $4.40 per 1M tokens. Same base as GLM-5.2; weights were still unpublished on this check. Ledger rows sync from AA; editorial depth still needs a desk pass. Compare it with Qwen3.8 27B. /models/glm-5-3 · /compare/models/glm-5-3-vs-qwen3-8-27b

    Drop
  6. Qwen3.8 27B is on the model atlas — AA index 52

    Alibaba's dense open-weight Qwen3.8 sibling is on the atlas from the 2026-08-24 Artificial Analysis snapshot: xhigh Intelligence Index 52, hosted list $0.50 / $3 per 1M tokens, Apache 2.0 on Hugging Face. This is not hosted Qwen3.8 Max. Ledger rows sync from AA; editorial depth still needs a desk pass. Compare it with Qwen3.8 Max. /models/qwen3-8-27b · /compare/models/qwen3-8-27b-vs-qwen3-8-max

    Drop

2026-W34

  1. GLM-5.3

    Hosted API listed on Artificial Analysis (max-effort index 59.5, $1.40 / $4.40 per 1M tokens).

    Model

2026-W33

  1. Qwen3.8 2.4T is on the model atlas — AA index 57.7

    Alibaba's open-weight Qwen-Max-class MoE (2.4T total / 95B active) is on the atlas from the 2026-08-15 Artificial Analysis snapshot: Intelligence Index 57.7, hosted list $2 / $6 per 1M tokens. This is the published checkpoint, not hosted Qwen3.8 Max. Ledger rows sync from AA; editorial depth still needs a desk pass. Compare it with Qwen3.8 Max. /models/qwen3-8-2-4t-a95b · /compare/models/qwen3-8-2-4t-a95b-vs-qwen3-8-max

    Drop
  2. GLM-5.3

    GLM-5.3 announced on the GLM Coding Plan. Open weights delayed for safety review.

    Model
  3. Qwen3.8 27B

    Open weights published. AA Intelligence Index 52 on xhigh; hosted list $0.50 / $3 per 1M tokens.

    Model
  4. Gemini 3.6 Flash (high)

    Objective refresh: DeepMind card 1M context / 64K output; Google intro price $0.75 / $3.75; AA high-effort index 51.6.

    Model
  5. Gemini 3.7 Flash

    Gemini 3.7 Flash generally available. AA high-effort Intelligence Index 56; intro price $0.75 / $3.75 per 1M tokens. Google reports gains over 3.6 Flash on FrontierCode 1.1, DeepSWE v1.1, WebDev Arena, and GDP.pdf — vendor-claimed, not ledger rows.

    Model
  6. Gemini 3.7 Flash is on the model atlas — AA index 56

    Google's August Flash workhorse is on the atlas from the Artificial Analysis snapshot: high-effort Intelligence Index 56, intro price $0.75 / $3.75 per 1M tokens through 31 Dec 2026. Ledger rows sync from AA; editorial depth still needs a desk pass. Compare it with 3.6 Flash or Claude Sonnet 4.6. /models/gemini-3-7-flash · /compare/models/gemini-3-6-flash-vs-gemini-3-7-flash

    Drop
  7. A.X-K2

    A.X-K2 appears on Artificial Analysis (index 35.0, $0 / $0).

    Model
  8. Grok 4.6

    Grok 4.6 appears on Artificial Analysis (high-effort index 60.9, $2 / $6 per 1M tokens).

    Model
  9. Grok 4.6 is on the model atlas — AA index 60.9

    xAI's August flagship is on the atlas from the Artificial Analysis snapshot: high-effort Intelligence Index 60.9, $2 / $6 per 1M tokens. Ledger rows sync from AA; editorial depth still needs a desk pass. Compare it with GPT-5.6 Sol. /models/grok-4-6 · /compare/models/gpt-5-6-sol-vs-grok-4-6

    Drop
  10. K-EXAONE 2.0 0803

    K-EXAONE 2.0 0803 appears on Artificial Analysis (index 31.0, $0 / $0).

    Model
  11. Motif 3

    Motif 3 GA on Artificial Analysis (index 47.4). Open-weight MoE, 314B / 13.2B active.

    Model
  12. Qwen3.8 2.4T A95B

    Open weights published. AA Intelligence Index 57.7; hosted list $2 / $6 per 1M tokens.

    Model
  13. Solar Open2 250B

    Solar Open2 250B appears on Artificial Analysis (index 37.4, $0 / $0).

    Model
  14. Nemotron 3.5 Lightning

    Nemotron 3.5 Lightning on Artificial Analysis (index 23.6, $0.05 / $0.20).

    Model
  15. Muse Glimmer

    Muse Glimmer open-weighted (Apache 2.0, 30B). AA high-effort index 35.1; hosted list $0.325 / $1.35 per 1M tokens.

    Model

2026-W32

  1. Solar Pro 4

    Solar Pro 4 on Artificial Analysis (index 41.6, $0.30 / $1.20).

    Model
  2. Muse Spark 1.2

    Muse Spark 1.2 appears on Artificial Analysis (xhigh index 56.8, $1.25 / $4.25 per 1M tokens).

    Model
  3. Qwen3.8 Max

    Qwen3.8 Max appears on Artificial Analysis (index 58.1, $2 / $6 per 1M tokens).

    Model
  4. VerdictPal is live, and the Atlas Report puts numbers on the whole atlas

    The launch note is published, plus a new original-research surface: the Atlas Report computes aggregate pricing, privacy, benchmark-saturation and freshness figures from every published tool, and it is free to cite with a link. Highlights this edition — most dossiers do not record any stance on training with your content, paid entry plans cluster around the median, and frontier model input prices spread by more than three orders of magnitude. /report · /launch

    Drop

2026-W29

  1. GPT-5.6 Luna

    Recorded general availability across ChatGPT, Codex, and the API and refreshed benchmark evidence.

    Model
  2. GPT-5.6 Sol

    Recorded general availability across ChatGPT, Codex, and the API, then refreshed Vals and Artificial Analysis benchmark evidence.

    Model
  3. GPT-5.6 Terra

    Recorded general availability across ChatGPT, Codex, and the API and refreshed benchmark evidence.

    Model

2026-W28

  1. AI Builder Design Test on the Lab bench — four builders, same prompts, very different results

    VerdictPal Lab's second first-party protocol: four identical multi-turn design prompts through Lovable, v0, Replit, and Bolt. Run 01 recorded a Turn 1 ranking (Lovable → v0 → Replit → Bolt) and the cost-of-arrival in each vendor's native unit — no padded leaderboard, formal scored rows pending a repeatable rubric. /benchmarks/ai-builder-design-test · /benchmarks/lab

    Drop
  2. Grok 4.5

    Grok 4.5 appears on Artificial Analysis (high-effort index 53.8).

    Model
  3. Citation Fidelity v0.1 protocol frozen on the Lab bench

    VerdictPal Lab's first first-party benchmark: twelve research questions, dual-reviewer citation grading, frozen runbook on the atlas page — protocol published, desk run still pending, no padded leaderboard. /benchmarks/citation-fidelity · /benchmarks/lab

    Drop
  4. Claude

    Evidence integrity fix: reverted false editorial-hands-on upgrade; evidence stays synthesized and qualityGate downgraded flagship → solid per the editorial evidence ladder. No desk hands-on trial on file; benchmarkReady remains false.

    Tool

2026-W27

  1. Five research guides now on the atlas

    Citation audit, systematic review protocol, data cleaning pipeline, grant writer path, and peer reviewer path — solid guides with step blocks and verification checklists.

    Drop
  2. GPT-5.6 Luna

    Updated public pricing: Luna is listed at $1 input / $6 output per 1M tokens.

    Model
  3. GPT-5.6 Sol

    Updated public pricing: Sol is listed at $5 input / $30 output per 1M tokens, with 1.25x cache-write billing and 90% cached-read discount.

    Model
  4. GPT-5.6 Terra

    Updated public pricing: Terra is listed at $2.50 input / $15 output per 1M tokens.

    Model
  5. OpenRouter

    Added LongCat-2.0/Owl Alpha release context: Meituan identified the OpenRouter stealth route as LongCat-backed after it had already become a high-volume coding model. Treat it as evidence of routing-market adoption, not as a privacy shortcut or independently verified benchmark row.

    Tool
  6. Claude Sonnet 5

    Claude Sonnet 5 listed on Artificial Analysis (max-effort index 53.4).

    Model
  7. LongCat-2.0

    Meituan revealed LongCat-2.0 and identified OpenRouter's Owl Alpha stealth route as LongCat-backed; GitHub/Hugging Face pages say model weights are coming soon.

    Model

2026-W26

  1. GPT-5.6 Luna

    Previewed as the fastest and most cost-efficient GPT-5.6 tier.

    Model
  2. GPT-5.6 Sol

    OpenAI previewed GPT-5.6 Sol alongside Terra and Luna.

    Model
  3. GPT-5.6 Terra

    Previewed as the balanced GPT-5.6 tier.

    Model
  4. Student research workflow on the guides pillar

    Path, stack, playbook, and showcase for turning broad prompts into source-backed briefs — the workflow we point new readers at first.

    Drop
  5. Factory flagship dossier shipped

    Agent-native delivery with Droid, Spec Mode, missions, and enterprise controls — fit 74 with the autonomy and pricing caveats on the record.

    Drop
  6. Cursor flagship dossier shipped

    VS Code-compatible AI editor at fit 81 — Tab, Chat, Agent, Privacy Mode, team governance, and the diff-review failure modes we publish before you merge.

    Drop
  7. Gemini flagship dossier shipped

    Google's multimodal assistant ecosystem at fit 75 — Search grounding, Deep Research, storage bundles, and where it is not a cited search engine.

    Drop

2026-W25

  1. NotebookLM flagship dossier shipped

    Source-grounded reading workspace at fit 78 — corpus chat, Audio Overviews, Flashcards, and the empty-corpus trap documented on the dossier.

    Drop
  2. Perplexity flagship dossier shipped

    Cited search front door at fit 72 — consumer Pro/Max, Deep Research, Sonar APIs, and failure modes for bibliographies you have not opened.

    Drop
  3. Cursor

    Production-readiness audit pass: editorsNote and pricing/privacy checked dates confirmed. benchmarkReady remains false intentionally until agent-diff safety, usage-burn, and Privacy Mode runs land.

    Tool
  4. Factory

    Production-readiness audit pass: editorsNote and pricing/privacy checked dates confirmed. benchmarkReady remains false intentionally until Spec Mode, Droid Exec autonomy, and enterprise-control runs land.

    Tool
  5. Gemini

    Production-readiness audit pass: editorsNote and pricing/privacy checked dates confirmed. No editorial claims changed; bench rows remain explicitly pending until the desk protocol runs (Next action in qualityGate.notes).

    Tool
  6. Perplexity

    Production-readiness audit pass: editorsNote, pricing/privacy checked dates confirmed; Sonar/Search API token + per-request fee rows aligned with docs.perplexity.ai. benchmarkReady remains false intentionally until desk citation audit and Sonar latency runs land.

    Tool
  7. Command Code

    Shipped Command Code flagship tool — terminal agent with taste learning, open-model deals, and full pricing/privacy research from commandcode.ai.

    Tool
  8. Command Code flagship dossier shipped

    Terminal coding agent with taste-1 learning, open-model deals from $1/mo, and the full pricing and privacy picture — including what lives in .commandcode/.

    Drop
  9. OpenCode

    Shipped OpenCode flagship tool — OSS agent, Go/Zen pricing, Plan/Build modes, and Zen privacy exceptions documented.

    Tool
  10. OpenCode flagship dossier shipped

    The open-source coding agent at fit 95 — Plan/Build modes, Go and Zen pricing, 75+ providers, and the free-model privacy caveats on the record.

    Drop
  11. Raycast

    Shipped Raycast as a flagship Knowledge work tool — the macOS launcher we would install on every research Mac before debating another AI tab.

    Tool
  12. Raycast flagship dossier shipped

    The macOS launcher we install before debating another AI tab — extensions, Quick AI, clipboard, and honest failure modes at fit 97.

    Drop

2026-W24

  1. GLM-5.2

    Released GLM-5.2 with 1M context, 128K max output, two thinking modes, and long-horizon coding focus. MIT open weights announced for the following week.

    Model
  2. Kimi K2.7 Code

    Released Kimi K2.7 Code with 256K context, mandatory thinking mode, multimodal tool examples, and reported 30% lower reasoning-token usage versus K2.6.

    Model
  3. Canva Business

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  4. ChatGPT

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  5. ChatPRD

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  6. Claude Fable 5

    Released Claude Fable 5 — first GA Mythos-class model.

    Model
  7. Connected Papers

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  8. Consensus

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  9. Crossref

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  10. DeepSeek

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  11. ElevenLabs

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  12. Elicit

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  13. Exa

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  14. Firecrawl

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  15. Gamma

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  16. Google Scholar

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  17. Grammarly

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  18. Gumloop

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  19. Kagi

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  20. Linear

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  21. Litmaps

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  22. Lovable

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  23. Magic Patterns

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  24. Manus

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  25. Mendeley

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  26. Microsoft Copilot

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  27. Mistral Le Chat

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  28. Mobbin

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  29. n8n

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  30. Obsidian

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  31. OpenAlex

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  32. OpenRouter

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  33. Overleaf

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  34. Paperpile

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  35. PostHog

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  36. Rayyan

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  37. Replit

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  38. Research Rabbit

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  39. Scholarcy

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  40. Scira AI

    Refreshed pricing evidence from scira.ai/about and api.scira.ai/pricing, added machine-readable pricing tiers, and marked benchmark evidence ready for the public dossier while keeping individual desk benchmark rows explicit about what still needs live runs.

    Tool
  41. Scite

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  42. Semantic Scholar

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  43. Warp

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  44. Wispr Flow

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  45. You.com

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  46. Zotero

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Tool
  47. Grant deadline set (archived)

    The grant-deadline path has been retired. Academic guides now live as three trios: student researcher, literature reviewer, and peer reviewer.

    Drop

2026-W23

  1. ChatGPT

    ChatGPT is the baseline general assistant tool, but it now needs to be framed as a family of surfaces: consumer ChatGPT, Business/Enterprise workspaces, Codex/agent features, and the separate OpenAI API. Its strength is breadth — drafting, file work, multimodal analysis, coding, agents, and extensions — while its risk is that fluent output can make weak sourcing, privacy assumptions, or plan limits invisible.

    Tool
  2. Claude

    Claude is the careful-drafting and long-context tool. The updated evidence reinforces its core lane: nuanced rewriting, document analysis, Projects, Artifacts, Research, Claude Code, Cowork, web search, and organization search/connectors. The editorial risk is quota and surface confusion: consumer Pro/Max, Team, Enterprise, and API/tool pricing are meaningfully different products.

    Tool
  3. Connected Papers

    Connected Papers remains the clean visual literature-map card. It is especially useful for orienting around one seed paper, but its core limitation is still methodological: a graph is not a reproducible search strategy. This pass adds official plan/feature detail, group/library plans, payment processing, scholarship support, and free-quota boundaries.

    Tool
  4. Consensus

    Consensus remains the claim-shaped academic Q&A tool. It is strongest when the user has a research question and needs cited orientation quickly. This pass adds current corpus/user-scale claims, Research Agent, account/privacy controls, institutional full-text linking, and a sharper “answer box is not a literature review” warning.

    Tool
  5. Crossref

    Crossref is scholarly infrastructure, not a shiny discovery app. It remains the DOI metadata verification card: use it for DOI records, member/publisher metadata, references, funder/license fields, relation metadata, and citation plumbing. This pass adds the 2025/2026 REST API rate-limit shift and the practical “polite pool” guidance that makes Crossref usable in responsible pipelines.

    Tool
  6. Cursor

    Cursor remains the AI-native code editor flagship, but the tool now needs to emphasize three separations: editor workflow versus headless agents, individual usage pools versus team/enterprise governance, and Privacy Mode versus ordinary provider transit. Cursor is excellent for repo-local Tab, Chat, Agent, MCP, cloud agents, and Bugbot workflows, but only if diffs are reviewed and tests run.

    Tool
  7. DeepSeek

    DeepSeek is the low-cost reasoning/API tool. This pass sharpens the key facts: DeepSeek’s API is OpenAI- and Anthropic-compatible, current V4-Flash/V4-Pro pricing is extremely low, context length is listed at 1M, and legacy deepseek-chat / deepseek-reasoner names are being deprecated. The low price is real, but the editorial warning is bigger than usual: privacy, residency, institutional policy, and geopolitical availability need explicit review.

    Tool
  8. ElevenLabs

    ElevenLabs is the voice and audio AI tool. This pass uses primary ElevenLabs pricing and privacy pages plus compliance/search results. The product’s quality can make outputs feel finished, but the editorial review must center voice consent, identity risk, disclosure, credit burn, and whether uploaded voices/interviews belong in a cloud voice system.

    Tool
  9. Elicit

    Elicit is no longer just a paper-search assistant; it is positioning itself as an evidence-synthesis system for reports, systematic reviews, dynamic screening, extraction, and sentence-level citations over a very large scientific corpus. The richer data supports a higher-confidence card, but the recommendation still depends on benchmarked extraction accuracy against a hand-coded table.

    Tool
  10. Exa

    Exa is one of the cleanest “search as infrastructure” cards in the VerdictPal stack. The current pricing and docs make the value clearer: real-time search, webpage text/highlights, configurable latency, contents retrieval, Answer endpoint, Deep Search, Monitors, and Agent runs. The core risk is economic and privacy-boundary creep when agents fan out across many searches, contents calls, summaries, and enrichment steps.

    Tool
  11. Factory

    Factory is the agent-native software-delivery tool. It is strongest when the task is bigger than autocomplete: repo understanding, spec-planned refactors, terminal/IDE delegation, background agents, SDK use, and CI-style Droid Exec workflows. This pass adds current 2026 Factory/Droid source detail and sharpens the enterprise/security boundary.

    Tool
  12. Firecrawl

    Firecrawl is the web-extraction infrastructure tool. It should be described as context extraction for agents and research pipelines, not as a generic search engine. This pass adds current credit economics, enterprise ZDR/SOC2 claims, endpoint credit costs, and the important warning that hosted scraping is always a trust, permission, and site-terms problem.

    Tool
  13. Gamma

    Gamma is the AI presentation and visual-document first-draft tool. This pass uses official Gamma pricing/product/privacy/search snippets, but the pricing and privacy pages timed out on full load, so the dossier clearly marks what is verified from retrieved official snippets and what needs final live-page verification before publication.

    Tool
  14. Gemini

    Deepened from official Google pages: pricing across consumer (AI Plus/Pro/Ultra), Workspace add-on, Gemini API paid tier, and Vertex AI; privacy language for consumer vs API vs Workspace/Vertex; capability list refreshed to current Gemini app + API surface.

    Tool
  15. Google Scholar

    Google Scholar remains the academic-search baseline: broad, familiar, fast, and useful for finding known items, PDFs, library links, citation trails, alerts, and case law. Its weakness is methodological opacity. It should be recommended as a discovery and citation-chasing habit, not as a reproducible search strategy by itself.

    Tool
  16. Grammarly

    Grammarly is the revision-pass tool, not the argument-writing tool. This pass updates the tool for Grammarly Pro and Superhuman-suite positioning: Free gives core correction plus 100 AI prompts, Pro adds rewrites/tone/brand features and 2,000 AI prompts, team features run up to 149 seats, and larger organizations move to Superhuman Go/enterprise-style controls. The core warning remains voice flattening and cloud text processing.

    Tool
  17. Gumloop

    Gumloop is the no-code AI workflow and agent-automation tool. This pass is based on primary Gumloop pricing and privacy pages, plus official product/search results. The review should focus on whether credit-based agent workflows stay understandable after real data, credentials, retries, and handoffs enter the loop.

    Tool
  18. Kagi

    Kagi is the paid-search tool: a user-funded, ad-free search engine with serious privacy engineering and enough Assistant functionality to be useful without turning into an AI slop layer. The strongest angle is incentive alignment — no ads, no result-click tracking, no analytics/telemetry on the main site, and paid plans that make the business model legible. The recommendation still depends on a 25-query search-quality benchmark against Google, Brave, DuckDuckGo, Perplexity, and You.com.

    Tool
  19. Litmaps

    Litmaps is the citation-map plus monitoring tool. It is valuable when a user starts from seed papers and needs to discover connected work, visualize the field, sync references, and keep watch for new papers. This pass fixes the malformed sources area and adds current Litmaps feature, coverage, and DPA details.

    Tool
  20. Mendeley

    Mendeley is the Elsevier-backed cloud-sync reference manager with increasingly AI-shaped library features. The richer data clarifies the buyer’s question: Mendeley is capable for PDF import, annotation, Word citation, groups, and AI over a library, but the deciding tradeoff is Elsevier cloud trust, storage/AI plan economics, and future switching cost versus Zotero and Paperpile.

    Tool
  21. Microsoft Copilot

    Microsoft Copilot is a family-name card, not a single-product card. This pass uses Microsoft Learn pages for Microsoft 365 Copilot and Copilot Chat data protection plus Microsoft pricing search results. The key editorial job is to separate consumer Copilot, Microsoft 365 Copilot Chat, paid Microsoft 365 Copilot, and Azure/OpenAI developer paths.

    Tool
  22. Mistral Le Chat

    Mistral Le Chat is the EU-headquartered assistant comparison tool. This pass strengthens the privacy and API boundary: Le Chat inputs/outputs are retained until account/conversation deletion, API inputs/outputs are generally kept for 30 rolling days for abuse monitoring unless zero data retention is activated, and Agents/Fine-tuning have separate retention rules. That makes Mistral useful to compare against US-default assistants, but not automatically compliant.

    Tool
  23. n8n

    n8n is the research-ops automation tool. It is strongest when a team needs visible workflow glue across APIs, webhooks, databases, AI steps, Notion, and alerts. The updated evidence makes the trade-off clearer: every plan now leans into unlimited users/workflows/steps, but pricing is based on full workflow executions, and self-hosted still means the operator owns patching, backups, secrets, telemetry choices, and failure recovery.

    Tool
  24. NotebookLM

    NotebookLM remains the source-grounded workspace flagship. The strongest VerdictPal framing is not “AI search,” but “corpus-first reading workspace”: upload or discover sources, ask grounded questions, generate study artifacts, and verify claims against citations. This pass adds current Enterprise, Workspace, and Google AI plan distinctions so the privacy story is clearer.

    Tool
  25. Obsidian

    Obsidian is the local-first knowledge-base tool. The important update is that Obsidian is now free for work as well as personal use, while paid services remain optional add-ons for Sync and Publish. That strengthens the tool’s editorial angle: Obsidian is not a cloud workspace trying to lock your notes away; it is a Markdown vault where the user owns the files, but also owns the backup, plugin, and structure decisions.

    Tool
  26. OpenAlex

    OpenAlex is the open scholarly graph tool. It sits between Crossref and Semantic Scholar: broader discovery and mapping than DOI registry metadata, more reproducible and open than closed search products, but still an aggregated graph that requires verification for final citations. This pass adds the 2026 usage-based pricing and privacy promise details.

    Tool
  27. OpenRouter

    OpenRouter is the model-router tool. It is useful when a developer wants one OpenAI-compatible API surface for many LLMs, fallback routing, model comparisons, BYOK, and spend controls. The updated tool sharpens the key caveat: routing convenience adds a platform fee and a subprocessor/provider-retention surface that must be controlled before confidential research or user data flows through it.

    Tool
  28. Overleaf

    Overleaf is the collaborative LaTeX tool. This pass updates the tool with current plan evidence, including AI allowance language, collaborator limits, 24x compile timeout on paid tiers, real-time track changes, and the subscriber-owned project entitlement model. The core trade remains: Overleaf removes local TeX and collaboration friction, but your project source and compile workflow live in a cloud editor.

    Tool
  29. Paperpile

    Paperpile is the convenience-first reference manager for Google Docs and browser-native writing. The stronger data now makes the tradeoff clearer: it has excellent Google Docs and PDF workflow coverage, annual pricing only, unlimited PDF storage on paid plans, and a cloud/Google-account privacy boundary that should be explicit for thesis and institutional work.

    Tool
  30. Perplexity

    Perplexity remains the cited-answer benchmark for consumer AI search, but the tool now needs to treat Perplexity as two related products: the consumer answer engine and the API platform. The consumer product is useful for fast source maps and Deep Research drafts; the API platform is a separate developer surface with Sonar, Search, Agent, and Embeddings APIs. The hard editorial rule stays the same: citations are leads, not bibliography-ready evidence.

    Tool
  31. PostHog

    PostHog is the product-analytics and experimentation-stack tool. It belongs in VerdictPal because product analytics, session replay, feature flags, experiments, surveys, pipelines, and AI/product ops determine how the product learns after launch. The recommendation stays gated because PostHog’s biggest risks are exactly the ones product teams underestimate: event volume, replay privacy, retention, and usage-based cost.

    Tool
  32. Rayyan

    Rayyan is a focused systematic-review screening workspace with stronger plan mechanics and AI boundaries than the earlier tool captured. The product’s editorial value is title/abstract screening, deduplication, PICO extraction, PRISMA support, reviewer/viewer collaboration, mobile work, and institutional ResearchPilot. The gold-standard warning is unchanged: AI can assist review logistics, but the final inclusion/exclusion decision must remain human and auditable.

    Tool
  33. Replit

    Replit is the browser-based AI build-and-deploy tool. It is useful when a user wants one place for ideation, agent building, database, hosting, collaboration, and publishing. The review must focus on code quality, credit burn, deployment security, and whether the user understands what the agent built.

    Tool
  34. Research Rabbit

    Research Rabbit is the visual exploration and “follow the trail” tool. It belongs beside Connected Papers and Litmaps, but its emphasis is adaptive exploration, collections, and seeing how papers/authors/concepts connect. This pass adds current feature language, data-source hints, and a clearer privacy/method boundary.

    Tool
  35. Scholarcy

    Scholarcy is the paper-skimming and structured-summary tool. It can save time for triage, accessibility, and organization, but it must be reviewed as a reading aid, not a reading replacement. This pass adds current product/pricing/help details, browser-extension behavior, export/library capabilities, and a sharper copyright/privacy boundary for uploaded PDFs.

    Tool
  36. Scira AI

    Scira remains the VerdictPal flagship template because it combines a cited search UI, open-source AGPL codebase, hosted Free/Pro/Max plans, an API platform, MCP surface, and a deep dossier structure. This pass replaces stale/malformed page content with a cleaner benchmarkable dossier while preserving the core stance: strong public recommendation, but benchmark-ready still false until the pending desk rows are actually run.

    Tool
  37. Scite

    Scite is the citation-context tool. It is valuable because it asks a better question than raw citation count: did later papers support, contrast, or merely mention the claim? This pass strengthens the tool with Scite’s current Smart Citations, MCP, API, publisher/full-text coverage, privacy-policy, and reference-check caveats. The public recommendation remains gated until label accuracy is benchmarked against actual citing passages.

    Tool
  38. Semantic Scholar

    Semantic Scholar remains a flagship academic-search layer because it combines a free scholarly search UI with TLDRs, author pages, alerts, a developer API, and downloadable scholarly graph data. It is more structured than Google Scholar and more paper-discovery oriented than Crossref. The warning is unchanged: AI summaries and graph edges are discovery aids, not evidence.

    Tool
  39. Warp

    Warp is the agentic terminal/workspace tool. It is compelling because it puts local and cloud coding agents where developers already run commands, but the evaluation must center command safety, repo hygiene, and data controls.

    Tool
  40. You.com

    You.com should now be reviewed primarily as a web-search API platform for AI builders, not as a nostalgic consumer search/chat competitor. Its current public pitch centers on Web Search APIs, Research API, zero-data-retention options, SOC2, DPA readiness, high rate limits, and enterprise deployment. The tool’s benchmark needs to test API source quality, freshness, cost, and data-control clarity against Exa, Perplexity Sonar/Search, Brave, Tavily-style APIs, and Kagi.

    Tool
  41. Zotero

    Zotero remains the strongest local-first default for student and academic reference management. The new data points reinforce the core positioning: the desktop workflow works without an account, sync is optional and disabled by default, paid storage funds a nonprofit project, and the main editorial risk is not capability but governance around sync, attachments, plugins, and group ownership.

    Tool
  42. Canva Business

    Imported from the Notion Tool pipeline (Card pipeline) with Scira-gate status corrected to solid.

    Tool
  43. ChatPRD

    Imported Notion state-of-art research and lowered the editorial fit to reflect weak desk usefulness plus Teams-gated Linear integration.

    Tool
  44. Linear

    Imported from the Notion Tool pipeline (Card pipeline) with Scira-gate status corrected to solid.

    Tool
  45. Lovable

    Imported from the Notion Tool pipeline (Card pipeline) with Scira-gate status corrected to solid.

    Tool
  46. Magic Patterns

    Imported from the Notion Tool pipeline (Card pipeline) with Scira-gate status corrected to solid.

    Tool
  47. Manus

    Imported Notion state-of-art fields and corrected Scira-gate status to solid.

    Tool
  48. Mobbin

    Imported from the Notion Tool pipeline (Card pipeline) with Scira-gate status corrected to solid.

    Tool
  49. Wispr Flow

    Imported from the Notion Tool pipeline (Card pipeline) with Scira-gate status corrected to solid.

    Tool
  50. Dark mode rebuilt for the evidence instrument

    Dark surfaces now use muted amber, readable borders, and component-specific overrides instead of neon panels and inverted cream panels.

    Drop
  51. Trusted-source brief workflow shipped

    The Guides pillar now has a real student research path, tool stack, playbook, and showcase for turning broad prompts into source-backed briefs.

    Drop
  52. Claude Mythos Preview

    Released Claude Mythos research preview — tops Vals (73.42%) and SWE-bench Verified (93.9%, best-of-3).

    Model

2026-W22

  1. MiniMax M3

    Released MiniMax M3 — Terminal-Bench 2.0 66.0%, 1M context.

    Model
  2. Claude Opus 4.8

    Released Claude Opus 4.8 — top of Vals (70.17%) and AA (61.4).

    Model

2026-W21

  1. Qwen 3.7 Max

    Released Qwen 3.7 Max — AA Index 56.6, Terminal-Bench 2.0 69.7%.

    Model
  2. Gemini 3.5 Flash

    Released Gemini 3.5 Flash — MCP Atlas leader (83.6%), 1M context, cheapest frontier pricing.

    Model

2026-W18

  1. Grok 4.3

    Released Grok 4.3 with 2M context and real-time X search.

    Model

2026-W17

  1. DeepSeek V4 Pro Max

    Released DeepSeek V4 Pro Max with open weights and frontier coding scores.

    Model
  2. GPT 5.5

    Released GPT 5.5 with native audio, image, and 512K context.

    Model
  3. GPT-5.5 Pro

    Released GPT-5.5 Pro — premium tier, $30/$180, Responses API.

    Model
  4. Kimi K2.6

    Released Kimi K2.6 — Vals #5 at 55.55%, GPQA Diamond 90.5% (open-source leader).

    Model

2026-W16

  1. Claude Opus 4.7

    Released Claude Opus 4.7 — 13% coding lift, 3x production tasks, 3.75 MP vision.

    Model
  2. Claude Opus 4.8

    Claude Opus 4.7 retired for general availability.

    Model

2026-W15

  1. GLM-5.1

    Released GLM-5.1 with 200K context, 128K max output, long-horizon coding focus, and MIT-licensed open weights.

    Model

2026-W14

  1. Qwen 3.7 Max

    Qwen 3.7 Plus retired for general availability.

    Model

2026-W12

  1. GPT-5.4 Mini

    Released GPT-5.4 Mini — best price/performance, $0.75/$4.50.

    Model
  2. GPT-5.4 Nano

    Released GPT-5.4 Nano — budget tier, $0.20/$1.25.

    Model

2026-W11

  1. MiniMax M3

    MiniMax M2.7 retired as the flagship open-weights model.

    Model

2026-W10

  1. GPT 5.4

    Released GPT 5.4 — AA Index 56.8, Terminal-Bench 2.0 81.8% (ForgeCode).

    Model
  2. GPT 5.5

    GPT 5.4 retired for general availability.

    Model
  3. GPT-5.4 Pro

    Released GPT-5.4 Pro — premium tier, $30/$180.

    Model
  4. Grok 4.3

    Grok 4.20 retired from general availability.

    Model

2026-W08

  1. Gemini 3.1 Pro

    Released Gemini 3.1 Pro — AA Index 57.2, ARC-AGI-2 77.1%.

    Model
  2. Claude Sonnet 4.6

    Released Claude Sonnet 4.6 — Vals #3 at 60.30%, 1M context, $3 / $15.

    Model

2026-W07

  1. DeepSeek V4 Pro Max

    DeepSeek V3.2 retired as the flagship open-weights model.

    Model

2026-W06

  1. Claude Opus 4.6

    Released Claude Opus 4.6 — first Opus with 1M context, 128K output, agent teams.

    Model

2025-W50

  1. GPT 5.4

    GPT 5.2 retired for general availability.

    Model
  2. GPT-5.2

    Released GPT-5.2 — 410K context, configurable reasoning, $1.75/$14.

    Model

2025-W48

  1. Claude Opus 4.5

    Released Claude Opus 4.5 — best coding and agent model at launch, $5/$25.

    Model

2025-W40

  1. Claude Haiku 4.5

    Released Claude Haiku 4.5 — speed tier, $1/$5, 200K context.

    Model
Pack drops feed