VerdictPal · Redaktions-Desk · 2026VerdictPal
Aktualität

Was sich änderte

Karten-Changelogs und Pack-Drops, nach Woche gruppiert.

2026-W29

  1. GPT-5.6 Luna

    Allgemeine Verfügbarkeit in ChatGPT, Codex und der API erfasst und Benchmark-Evidenz aktualisiert.

    Modell
  2. GPT-5.6 Sol

    Allgemeine Verfügbarkeit in ChatGPT, Codex und der API erfasst und anschließend Benchmark-Evidenz von Vals und Artificial Analysis aktualisiert.

    Modell
  3. GPT-5.6 Terra

    Allgemeine Verfügbarkeit in ChatGPT, Codex und der API erfasst und Benchmark-Evidenz aktualisiert.

    Modell

2026-W28

  1. AI-Builder-Design-Test auf der Lab-Bank — vier Builder, gleiche Prompts, sehr verschiedene Ergebnisse

    VerdictPals zweites First-Party-Protokoll: vier identische Multi-Turn-Design-Prompts durch Lovable, v0, Replit und Bolt. Lauf 01 notiert eine Turn-1-Rangfolge (Lovable → v0 → Replit → Bolt) und die Kosten bis zum Ergebnis in der jeweiligen Hersteller-Einheit — kein aufgefülltes Leaderboard, formale Score-Zeilen warten auf ein wiederholbares Rubric. /benchmarks/ai-builder-design-test · /benchmarks/lab

    Drop
  2. Grok 4.5

    Grok 4.5 erscheint auf Artificial Analysis (High-Effort-Index 53,8).

    Modell
  3. Citation-Fidelity-Protokoll v0.1 auf der Lab-Bank eingefroren

    VerdictPals erster First-Party-Benchmark im Lab: zwölf Forschungsfragen, dual-reviewer Zitierbewertung, eingefrorenes Runbook auf der Atlas-Karte — Protokoll veröffentlicht, Desk-Lauf noch ausstehend, kein aufgefülltes Leaderboard. /benchmarks/citation-fidelity · /benchmarks/lab

    Drop
  4. Claude

    Evidence integrity fix: reverted false editorial-hands-on upgrade; evidence stays synthesized and qualityGate downgraded flagship → solid per docs/EDITORIAL-SOP.md §2 evidence ladder. No desk hands-on trial on file; benchmarkReady remains false.

    Karte

2026-W27

  1. Fünf Recherche-Guides jetzt im Atlas

    Zitier-Audit, Systematic-Review-Protokoll, Data-Cleaning-Pipeline, Grant-Writer-Pfad und Peer-Reviewer-Pfad — solide Guides mit Step-Cards und Verifikations-Checklisten.

    Drop
  2. GPT-5.6 Luna

    Öffentliche Preise aktualisiert: Luna ist mit $1 Input / $6 Output pro 1M Tokens gelistet.

    Modell
  3. GPT-5.6 Sol

    Öffentliche Preise aktualisiert: Sol ist mit $5 Input / $30 Output pro 1M Tokens gelistet, plus 1,25× Cache-Write-Billing und 90% Cached-Read-Rabatt.

    Modell
  4. GPT-5.6 Terra

    Öffentliche Preise aktualisiert: Terra ist mit $2,50 Input / $15 Output pro 1M Tokens gelistet.

    Modell
  5. OpenRouter

    LongCat-2.0-/Owl-Alpha-Release-Kontext ergänzt: Meituan identifizierte die OpenRouter-Stealth-Route als LongCat-gestützt, nachdem sie bereits ein volumenstarkes Coding-Modell war. Als Evidenz für Routing-Markt-Adoption behandeln, nicht als Datenschutz-Abkürzung oder unabhängig verifizierte Benchmark-Zeile.

    Karte
  6. Claude Sonnet 5

    Claude Sonnet 5 auf Artificial Analysis gelistet (Max-Effort-Index 53,4).

    Modell
  7. LongCat-2.0

    Meituan enthüllte LongCat-2.0 und identifizierte OpenRouters Stealth-Route Owl Alpha als LongCat-gestützt; GitHub/Hugging-Face-Seiten sagen, dass Modellgewichte bald folgen.

    Modell

2026-W26

  1. GPT-5.6 Luna

    Als schnellster und kosteneffizientester GPT-5.6-Tier vorgestellt.

    Modell
  2. GPT-5.6 Sol

    OpenAI hat GPT-5.6 Sol zusammen mit Terra und Luna als Preview vorgestellt.

    Modell
  3. GPT-5.6 Terra

    Als ausgewogener GPT-5.6-Tier vorgestellt.

    Modell
  4. Studenten-Recherche-Workflow in der Guides-Säule

    Pfad, Stack, Playbook und Showcase, um breite Prompts in quellenbasierte Briefings zu verwandeln — der Workflow, den wir neuen Leserinnen zuerst zeigen.

    Drop
  5. Factory-Flaggschiff-Dossier veröffentlicht

    Agenten-native Auslieferung mit Droid, Spec Mode, Missionen und Enterprise-Kontrollen — Fit 74 mit dokumentierten Autonomie- und Preis-Ausnahmen.

    Drop
  6. Cursor-Flaggschiff-Dossier veröffentlicht

    VS-Code-kompatibler KI-Editor bei Fit 81 — Tab, Chat, Agent, Privacy Mode, Team-Governance und Diff-Review-Fehlermodi, die wir vor dem Merge dokumentieren.

    Drop
  7. Gemini-Flaggschiff-Dossier veröffentlicht

    Googles multimodales Assistenten-Ökosystem bei Fit 75 — Suchfundierung, Deep Research, Speicher-Bundles und wo es keine zitierte Suchmaschine ist.

    Drop

2026-W25

  1. NotebookLM-Flaggschiff-Dossier veröffentlicht

    Quellengestützter Lese-Workspace bei Fit 78 — Korpus-Chat, Audio-Übersichten, Lernkarten und die leere-Korpus-Falle auf der Karte dokumentiert.

    Drop
  2. Perplexity-Flaggschiff-Dossier veröffentlicht

    Zitierte Such-Eingangstür bei Fit 72 — Consumer Pro/Max, Deep Research, Sonar-APIs und Fehlermodi für Bibliografien, die du nicht geöffnet hast.

    Drop
  3. Cursor

    Production-readiness audit pass: editorsNote and pricing/privacy checked dates confirmed. benchmarkReady remains false intentionally until agent-diff safety, usage-burn, and Privacy Mode runs land.

    Karte
  4. Factory

    Production-readiness audit pass: editorsNote and pricing/privacy checked dates confirmed. benchmarkReady remains false intentionally until Spec Mode, Droid Exec autonomy, and enterprise-control runs land.

    Karte
  5. Gemini

    Production-readiness audit pass: editorsNote and pricing/privacy checked dates confirmed. No editorial claims changed; bench rows remain explicitly pending until the desk protocol runs (Next action in qualityGate.notes).

    Karte
  6. Perplexity

    Production-readiness audit pass: editorsNote, pricing/privacy checked dates confirmed; Sonar/Search API token + per-request fee rows aligned with docs.perplexity.ai. benchmarkReady remains false intentionally until desk citation audit and Sonar latency runs land.

    Karte
  7. Command Code

    Command-Code-Flaggschiff-Karte veröffentlicht — Terminal-Agent mit Taste-Lernen, Open-Model-Deals und vollständiger Preis-/Datenschutz-Recherche.

    Karte
  8. Command-Code-Flaggschiff-Dossier veröffentlicht

    Terminal-Coding-Agent mit taste-1-Lernen, Open-Model-Deals ab 1 $/Monat und vollständigem Preis- und Datenschutzbild — inklusive .commandcode/.

    Drop
  9. OpenCode

    OpenCode-Flaggschiff-Karte — OSS-Agent, Go/Zen-Preise, Plan-/Build-Modi und Zen-Datenschutz-Ausnahmen dokumentiert.

    Karte
  10. OpenCode-Flaggschiff-Dossier veröffentlicht

    Der Open-Source-Coding-Agent bei Fit 95 — Plan-/Build-Modi, Go- und Zen-Preise, 75+ Anbieter und Free-Model-Datenschutz-Ausnahmen dokumentiert.

    Drop
  11. Raycast

    Raycast als Flaggschiff-Karte für Knowledge work — der macOS-Launcher, den wir auf jedem Forschungs-Mac installieren würden, bevor wir über einen weiteren KI-Tab debattieren.

    Karte
  12. Raycast-Flaggschiff-Dossier veröffentlicht

    Der macOS-Launcher, den wir installieren, bevor wir über einen weiteren KI-Tab debattieren — Extensions, Quick AI, Zwischenablage und ehrliche Fehlermodi bei Fit 97.

    Drop

2026-W24

  1. GLM-5.2

    GLM-5.2 mit 1M-Kontext, 128K Max-Output, zwei Thinking-Modi und Long-Horizon-Coding-Fokus veröffentlicht. MIT-Open-Weights für die folgende Woche angekündigt.

    Modell
  2. Kimi K2.7 Code

    Kimi K2.7 Code mit 256K-Kontext, verpflichtendem Thinking-Modus, multimodalen Tool-Beispielen und berichteten 30 % weniger Reasoning-Token gegenüber K2.6 veröffentlicht.

    Modell
  3. Canva Business

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  4. ChatGPT

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  5. ChatPRD

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  6. Claude Fable 5

    Claude Fable 5 veröffentlicht — erstes GA Mythos-class-Modell; Vals-Index 75,15 %.

    Modell
  7. Connected Papers

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  8. Consensus

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  9. Crossref

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  10. DeepSeek

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  11. ElevenLabs

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  12. Elicit

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  13. Exa

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  14. Firecrawl

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  15. Gamma

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  16. Google Scholar

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  17. Grammarly

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  18. Gumloop

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  19. Kagi

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  20. Linear

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  21. Litmaps

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  22. Lovable

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  23. Magic Patterns

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  24. Manus

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  25. Mendeley

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  26. Microsoft Copilot

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  27. Mistral Le Chat

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  28. Mobbin

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  29. n8n

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  30. Obsidian

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  31. OpenAlex

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  32. OpenRouter

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  33. Overleaf

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  34. Paperpile

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  35. PostHog

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  36. Rayyan

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  37. Replit

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  38. Research Rabbit

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  39. Scholarcy

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  40. Scira AI

    Refreshed pricing evidence from scira.ai/about and api.scira.ai/pricing, added machine-readable pricing tiers, and marked benchmark evidence ready for the public dossier while keeping individual desk benchmark rows explicit about what still needs live runs.

    Karte
  41. Scite

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  42. Semantic Scholar

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  43. Warp

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  44. Wispr Flow

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  45. You.com

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  46. Zotero

    Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.

    Karte
  47. Förderantrag-Set veröffentlicht

    Ein Flagship-Pfad, Stack, Playbook und Walkthrough — von der leeren Aims-Seite zu einreichungsfähigem One-Pager und Budgetbegründung in unter drei Stunden.

    Drop

2026-W23

  1. ChatGPT

    ChatGPT ist das/die baseline general assistant card, but it now needs to be framed as a family of surfaces: consumer ChatGPT, Business/Enterprise workspaces, Codex/agent features, and the separate OpenAI API. Its strength is breadth — drafting, file work, multimodal analysis, coding, agents, and extensions — while its risk is that fluent output can make weak sourcing, privacy assumptions, or plan limits invisible.

    Karte
  2. Claude

    Claude is the careful-drafting and long-context card. The updated evidence reinforces its core lane: nuanced rewriting, document analysis, Projects, Artifacts, Research, Claude Code, Cowork, web search, and organization search/connectors. The editorial risk is quota and surface confusion: consumer Pro/Max, Team, Enterprise, and API/tool pricing are meaningfully different products.

    Karte
  3. Connected Papers

    Connected Papers bleibt die saubere visuelle Literature-Map-Karte. Besonders nützlich zur Orientierung um ein Seed-Paper, aber die Kernlimitation bleibt methodisch: ein Graph ist kein reproduzierbarer Suchkorpus.

    Karte
  4. Consensus

    Consensus bleibt die behauptungsförmige akademische Q&A-Karte. Am stärksten, wenn der Nutzer eine Forschungsfrage hat und schnell zitierte Orientierung braucht.

    Karte
  5. Crossref

    Crossref ist wissenschaftliche Infrastruktur, keine glänzende Discovery-App. Es bleibt die DOI-Metadaten-Verifikationskarte.

    Karte
  6. Cursor

    Cursor remains the AI-native code editor flagship, but the card now needs to emphasize three separations: editor workflow versus headless agents, individual usage pools versus team/enterprise governance, and Privacy Mode versus ordinary provider transit. Cursor is excellent for repo-local Tab, Chat, Agent, MCP, cloud agents, and Bugbot workflows, but only if diffs are reviewed and tests run.

    Karte
  7. DeepSeek

    DeepSeek is the low-cost reasoning/API card. This pass sharpens the key facts: DeepSeek’s API is OpenAI- and Anthropic-compatible, current V4-Flash/V4-Pro pricing is extremely low, context length is listed at 1M, and legacy deepseek-chat / deepseek-reasoner names are being deprecated. The low price is real, but the editorial warning is bigger than usual: privacy, residency, institutional policy, and geopolitical availability need explicit review.

    Karte
  8. ElevenLabs

    ElevenLabs ist die Sprach- und Audio-KI-Karte. Die Produktqualität kann Ergebnisse ausgereift erscheinen lassen.

    Karte
  9. Elicit

    Elicit ist nicht länger nur ein Paper-Such-Assistent; es positioniert sich als Evidenzsynthese-System für Berichte, systematische Reviews und satzgenaue Zitationen.

    Karte
  10. Exa

    Exa is one of the cleanest “search as infrastructure” cards in the VerdictPal stack. The current pricing and docs make the value clearer: real-time search, webpage text/highlights, configurable latency, contents retrieval, Answer endpoint, Deep Search, Monitors, and Agent runs. The core risk is economic and privacy-boundary creep when agents fan out across many searches, contents calls, summaries, and enrichment steps.

    Karte
  11. Factory

    Factory is the agent-native software-delivery card. It is strongest when the task is bigger than autocomplete: repo understanding, spec-planned refactors, terminal/IDE delegation, background agents, SDK use, and CI-style Droid Exec workflows. This pass adds current 2026 Factory/Droid source detail and sharpens the enterprise/security boundary.

    Karte
  12. Firecrawl

    Firecrawl ist die Web-Extraktions-Infrastrukturkarte. Sollte als Kontext-Extraktion für Agenten und Forschungspipelines beschrieben werden, nicht als generische Suchmaschine.

    Karte
  13. Gamma

    Gamma ist die KI-Präsentations- und Visual-Document-Erstdraft-Karte. Die Preis- und Datenschutzseiten waren beim Laden nicht erreichbar.

    Karte
  14. Gemini

    Deepened from official Google pages: pricing across consumer (AI Plus/Pro/Ultra), Workspace add-on, Gemini API paid tier, and Vertex AI; privacy language for consumer vs API vs Workspace/Vertex; capability list refreshed to current Gemini app + API surface.

    Karte
  15. Google Scholar

    Google Scholar bleibt die akademische Such-Baseline: breit, vertraut, schnell und nützlich zum Finden bekannter Einträge, PDFs, Bibliothekslinks, Zitationsspuren, Alerts und Rechtsprechung.

    Karte
  16. Grammarly

    Grammarly ist die Überarbeitungsdurchlauf-Karte, nicht die Argument-Schreib-Karte. Dieser Durchlauf aktualisiert die Karte für Grammarly Pro und Superhuman-Suite.

    Karte
  17. Gumloop

    Gumloop ist die No-Code-KI-Workflow- und Agent-Automation-Karte. Basiert auf primären Preis- und Datenschutzseiten plus offiziellen Produkt-/Suchergebnissen.

    Karte
  18. Kagi

    Kagi ist die Paid-Search-Karte: eine nutzerfinanzierte, werbefreie Suchmaschine mit ernsthafter Datenschutz-Technik und nützlicher Assistant-Funktionalität.

    Karte
  19. Litmaps

    Litmaps ist die Citation-Map-plus-Monitoring-Karte. Wertvoll, wenn Nutzer von Seed-Papern ausgehen und vernetzte Arbeiten entdecken, visualisieren und überwachen wollen.

    Karte
  20. Mendeley

    Mendeley ist der Elsevier-gestützte Cloud-Sync-Referenzmanager mit zunehmend KI-geprägten Bibliotheksfunktionen.

    Karte
  21. Microsoft Copilot

    Microsoft Copilot ist a family-name card, not a single-product card. This pass uses Microsoft Learn pages for Microsoft 365 Copilot and Copilot Chat data protection plus Microsoft pricing search results. The key editorial job is to separate consumer Copilot, Microsoft 365 Copilot Chat, paid Microsoft 365 Copilot, and Azure/OpenAI developer paths.

    Karte
  22. Mistral Le Chat

    Mistral Le Chat ist is the EU-headquartered assistant comparison card. This pass strengthens the privacy and API boundary: Le Chat inputs/outputs are retained until account/conversation deletion, API inputs/outputs are generally kept for 30 rolling days for abuse monitoring unless zero data retention is activated, and Agents/Fine-tuning have separate retention rules. That makes Mistral useful to compare against US-default assistants, but not automatically compliant.

    Karte
  23. n8n

    n8n ist die Research-Ops-Automatisierungskarte. Am stärksten, wenn ein Team sichtbaren Workflow-Kleber über APIs, Webhooks, Datenbanken, KI-Schritte, Notion und Alerts benötigt.

    Karte
  24. NotebookLM

    NotebookLM remains the source-grounded workspace flagship. The strongest VerdictPal framing is not “AI search,” but “corpus-first reading workspace”: upload or discover sources, ask grounded questions, generate study artifacts, and verify claims against citations. This pass adds current Enterprise, Workspace, and Google AI plan distinctions so the privacy story is clearer.

    Karte
  25. Obsidian

    Obsidian ist die Local-First-Wissensdatenbank-Karte. Das wichtige Update: Obsidian ist jetzt sowohl für die Arbeit als auch privat kostenlos; Sync und Publish bleiben optional.

    Karte
  26. OpenAlex

    OpenAlex ist die offene wissenschaftliche Graph-Karte. Sie sitzt zwischen Crossref und Semantic Scholar: breitere Discovery als DOI-Registry, reproduzierbarer als geschlossene Suchprodukte.

    Karte
  27. OpenRouter

    OpenRouter ist die Model-Router-Karte. Nützlich, wenn ein Entwickler eine OpenAI-kompatible API-Oberfläche für viele LLMs, Fallback-Routing, Modellvergleiche und BYOK möchte.

    Karte
  28. Overleaf

    Overleaf ist die kollaborative LaTeX-Karte. Dieser Durchlauf aktualisiert die Karte mit aktuellen Plan-Nachweisen: KI-Kontingent, Mitarbeiterlimits und 24x-Kompilierzeit.

    Karte
  29. Paperpile

    Paperpile ist der Convenience-First-Referenzmanager für Google Docs und browser-natives Schreiben. Exzellente Google-Docs- und PDF-Workflow-Abdeckung, aber Tradeoffs bei bibliographischer Tiefe.

    Karte
  30. Perplexity

    Perplexity remains the cited-answer benchmark for consumer AI search, but the card now needs to treat Perplexity as two related products: the consumer answer engine and the API platform. The consumer product is useful for fast source maps and Deep Research drafts; the API platform is a separate developer surface with Sonar, Search, Agent, and Embeddings APIs. The hard editorial rule stays the same: citations are leads, not bibliography-ready evidence.

    Karte
  31. PostHog

    PostHog ist die Produktanalyse- und Experimentier-Stack-Karte. Sie gehört in VerdictPal, weil Produktanalyse, Session Replay, Feature Flags und Experimente sich mit Forschungs- und Produkt-Ops überschneiden.

    Karte
  32. Rayyan

    Rayyan ist ein fokussierter Systematic-Review-Screening-Workspace mit stärkeren Plan-Mechaniken.

    Karte
  33. Replit

    Replit ist die browserbasierte KI-Build-and-Deploy-Karte. Nützlich, wenn ein Nutzer einen Ort für Ideenfindung, Agent-Building, Datenbank, Hosting, Zusammenarbeit und Publishing möchte.

    Karte
  34. Research Rabbit

    Research Rabbit ist the visual exploration and “follow the trail” card. It belongs beside Connected Papers and Litmaps, but its emphasis is adaptive exploration, collections, and seeing how papers/authors/concepts connect. This pass adds current feature language, data-source hints, and a clearer privacy/method boundary.

    Karte
  35. Scholarcy

    Scholarcy ist die Paper-Überflug- und strukturierte-Zusammenfassungs-Karte. Kann Zeit für Triage, Accessibility und Organisation sparen, muss aber als Lesehilfe geprüft werden.

    Karte
  36. Scira AI

    Scira remains the VerdictPal flagship template because it combines a cited search UI, open-source AGPL codebase, hosted Free/Pro/Max plans, an API platform, MCP surface, and a deep dossier structure. This pass replaces stale/malformed page content with a cleaner benchmarkable dossier while preserving the core stance: strong public recommendation, but benchmark-ready still false until the pending desk rows are actually run.

    Karte
  37. Scite

    Scite ist die Zitationskontext-Karte. Wertvoll, weil sie eine bessere Frage stellt: Haben spätere Paper die Behauptung gestützt, kontrastiert oder nur erwähnt?

    Karte
  38. Semantic Scholar

    Semantic Scholar bleibt eine Flagship-akademische-Suchschicht, weil sie eine kostenlose wissenschaftliche Such-UI mit TLDRs, Autorenseiten, Alerts und Entwickler-API kombiniert.

    Karte
  39. Warp

    Warp ist die agentische Terminal/Workspace-Karte. Überzeugend, weil sie Coding-Agenten ins Terminal bringt.

    Karte
  40. You.com

    You.com sollte primär als Web-Search-API-Plattform für KI-Entwickler betrachtet werden.

    Karte
  41. Zotero

    Zotero bleibt der stärkste Local-First-Standard für studentisches und akademisches Referenzmanagement.

    Karte
  42. Canva Business

    Aus der Notion-Karten-Pipeline importiert mit Scira-gate status corrected to solid.

    Karte
  43. ChatPRD

    Notion-State-of-Art-Recherche importiert und redaktionelle Eignung gesenkt, um schwache Desk-Nützlichkeit plus Teams-gesicherte Linear-Integration widerzuspiegeln.

    Karte
  44. Linear

    Aus der Notion-Karten-Pipeline importiert mit Scira-gate status corrected to solid.

    Karte
  45. Lovable

    Aus der Notion-Karten-Pipeline importiert mit Scira-gate status corrected to solid.

    Karte
  46. Magic Patterns

    Aus der Notion-Karten-Pipeline importiert mit Scira-gate status corrected to solid.

    Karte
  47. Manus

    Notion-State-of-Art-Felder importiert und Scira-Gate-Status auf Solid korrigiert.

    Karte
  48. Mobbin

    Aus der Notion-Karten-Pipeline importiert mit Scira-gate status corrected to solid.

    Karte
  49. Wispr Flow

    Aus der Notion-Karten-Pipeline importiert mit Scira-gate status corrected to solid.

    Karte
  50. Dark Mode für das Evidence Instrument überarbeitet

    Dunkle Oberflächen nutzen jetzt gedämpftes Amber, lesbare Ränder und komponentenspezifische Overrides statt Neon-Karten und invertierter Cream-Panels.

    Drop
  51. Trusted-Source-Briefing-Workflow veröffentlicht

    Die Guides-Säule hat jetzt einen echten Recherchepfad, Tool-Stack, Playbook und Showcase, um breite Prompts in quellenbasierte Briefings zu verwandeln.

    Drop
  52. Claude Mythos Preview

    Claude-Mythos-Research-Preview veröffentlicht — Spitze bei Vals (73,42 %) und SWE-bench Verified (93,9 %, Best-of-3).

    Modell

2026-W22

  1. MiniMax M3

    MiniMax M3 veröffentlicht — Terminal-Bench 2.0 66,0 %, 1M-Kontext.

    Modell
  2. Claude Opus 4.8

    Claude Opus 4.8 veröffentlicht — Spitze bei Vals (70,17 %) und AA (61,4).

    Modell

2026-W21

  1. Qwen 3.7 Max

    Qwen 3.7 Max veröffentlicht — AA-Index 56,6, Terminal-Bench 2.0 69,7 %.

    Modell
  2. Gemini 3.5 Flash

    Gemini 3.5 Flash veröffentlicht — MCP-Atlas-Spitze (83,6 %), 1M-Kontext, günstigster Frontier-Preis.

    Modell

2026-W18

  1. Grok 4.3

    Grok 4.3 mit 2M-Kontext und Echtzeit-X-Suche veröffentlicht.

    Modell

2026-W17

  1. DeepSeek V4 Pro Max

    DeepSeek V4 Pro Max mit Open Weights und Frontier-Coding-Scores veröffentlicht.

    Modell
  2. GPT 5.5

    GPT 5.5 mit nativem Audio, Bild und 512K-Kontext veröffentlicht.

    Modell
  3. GPT-5.5 Pro

    GPT-5.5 Pro veröffentlicht — Premium-Stufe, $30/$180, Responses API.

    Modell
  4. Kimi K2.6

    Kimi K2.6 veröffentlicht — Vals #5 bei 55,55 %, GPQA Diamond 90,5 % (Open-Source-Spitze).

    Modell

2026-W16

  1. Claude Opus 4.7

    Claude Opus 4.7 veröffentlicht — 13 % Coding-Steigerung, 3× Produktionsaufgaben, 3,75 MP Vision.

    Modell
  2. Claude Opus 4.8

    Claude Opus 4.7 für allgemeine Verfügbarkeit ausgemustert.

    Modell

2026-W15

  1. GLM-5.1

    GLM-5.1 mit 200K-Kontext, 128K Max-Output, Long-Horizon-Coding-Fokus und MIT-lizenzierten Open Weights veröffentlicht.

    Modell

2026-W14

  1. Qwen 3.7 Max

    Qwen 3.7 Plus für allgemeine Verfügbarkeit ausgemustert.

    Modell

2026-W12

  1. GPT-5.4 Mini

    GPT-5.4 Mini veröffentlicht — bestes Preis-/Leistungs-Verhältnis, $0,75/$4,50.

    Modell
  2. GPT-5.4 Nano

    GPT-5.4 Nano veröffentlicht — Budget-Stufe, $0,20/$1,25.

    Modell

2026-W11

  1. MiniMax M3

    MiniMax M2.7 als Flaggschiff-Open-Weights-Modell ausgemustert.

    Modell

2026-W10

  1. GPT 5.4

    GPT 5.4 veröffentlicht — AA-Index 56,8, Terminal-Bench 2.0 81,8 % (ForgeCode).

    Modell
  2. GPT 5.5

    GPT 5.4 für allgemeine Verfügbarkeit ausgemustert.

    Modell
  3. GPT-5.4 Pro

    GPT-5.4 Pro veröffentlicht — Premium-Stufe, $30/$180.

    Modell
  4. Grok 4.3

    Grok 4.20 aus der allgemeinen Verfügbarkeit ausgemustert.

    Modell

2026-W08

  1. Gemini 3.1 Pro

    Gemini 3.1 Pro veröffentlicht — AA-Index 57,2, ARC-AGI-2 77,1 %.

    Modell
  2. Claude Sonnet 4.6

    Claude Sonnet 4.6 veröffentlicht — Vals #3 bei 60,30 %, 1M-Kontext, $3 / $15.

    Modell

2026-W07

  1. DeepSeek V4 Pro Max

    DeepSeek V3.2 als Flaggschiff-Open-Weights-Modell ausgemustert.

    Modell

2026-W06

  1. Claude Opus 4.6

    Claude Opus 4.6 veröffentlicht — erstes Opus mit 1M-Kontext, 128K-Output, Agent-Teams.

    Modell

2025-W50

  1. GPT 5.4

    GPT 5.2 für allgemeine Verfügbarkeit ausgemustert.

    Modell
  2. GPT-5.2

    GPT-5.2 veröffentlicht — 410K-Kontext, konfigurierbares Reasoning, $1,75/$14.

    Modell

2025-W48

  1. Claude Opus 4.5

    Claude Opus 4.5 veröffentlicht — bestes Coding- und Agent-Modell bei Launch, $5/$25.

    Modell

2025-W40

  1. Claude Haiku 4.5

    Claude Haiku 4.5 veröffentlicht — Speed-Stufe, $1/$5, 200K-Kontext.

    Modell
Pack-Drops-Feed