Was sich änderte
2026-W29
- GPT-5.6 Luna
Allgemeine Verfügbarkeit in ChatGPT, Codex und der API erfasst und Benchmark-Evidenz aktualisiert.
Modell - GPT-5.6 Sol
Allgemeine Verfügbarkeit in ChatGPT, Codex und der API erfasst und anschließend Benchmark-Evidenz von Vals und Artificial Analysis aktualisiert.
Modell - GPT-5.6 Terra
Allgemeine Verfügbarkeit in ChatGPT, Codex und der API erfasst und Benchmark-Evidenz aktualisiert.
Modell
2026-W28
- AI-Builder-Design-Test auf der Lab-Bank — vier Builder, gleiche Prompts, sehr verschiedene Ergebnisse
VerdictPals zweites First-Party-Protokoll: vier identische Multi-Turn-Design-Prompts durch Lovable, v0, Replit und Bolt. Lauf 01 notiert eine Turn-1-Rangfolge (Lovable → v0 → Replit → Bolt) und die Kosten bis zum Ergebnis in der jeweiligen Hersteller-Einheit — kein aufgefülltes Leaderboard, formale Score-Zeilen warten auf ein wiederholbares Rubric. /benchmarks/ai-builder-design-test · /benchmarks/lab
Drop - Citation-Fidelity-Protokoll v0.1 auf der Lab-Bank eingefroren
VerdictPals erster First-Party-Benchmark im Lab: zwölf Forschungsfragen, dual-reviewer Zitierbewertung, eingefrorenes Runbook auf der Atlas-Karte — Protokoll veröffentlicht, Desk-Lauf noch ausstehend, kein aufgefülltes Leaderboard. /benchmarks/citation-fidelity · /benchmarks/lab
Drop - Claude
Evidence integrity fix: reverted false editorial-hands-on upgrade; evidence stays synthesized and qualityGate downgraded flagship → solid per docs/EDITORIAL-SOP.md §2 evidence ladder. No desk hands-on trial on file; benchmarkReady remains false.
Karte
2026-W27
- Fünf Recherche-Guides jetzt im Atlas
Zitier-Audit, Systematic-Review-Protokoll, Data-Cleaning-Pipeline, Grant-Writer-Pfad und Peer-Reviewer-Pfad — solide Guides mit Step-Cards und Verifikations-Checklisten.
Drop - GPT-5.6 Luna
Öffentliche Preise aktualisiert: Luna ist mit $1 Input / $6 Output pro 1M Tokens gelistet.
Modell - GPT-5.6 Sol
Öffentliche Preise aktualisiert: Sol ist mit $5 Input / $30 Output pro 1M Tokens gelistet, plus 1,25× Cache-Write-Billing und 90% Cached-Read-Rabatt.
Modell - GPT-5.6 Terra
Öffentliche Preise aktualisiert: Terra ist mit $2,50 Input / $15 Output pro 1M Tokens gelistet.
Modell - OpenRouter
LongCat-2.0-/Owl-Alpha-Release-Kontext ergänzt: Meituan identifizierte die OpenRouter-Stealth-Route als LongCat-gestützt, nachdem sie bereits ein volumenstarkes Coding-Modell war. Als Evidenz für Routing-Markt-Adoption behandeln, nicht als Datenschutz-Abkürzung oder unabhängig verifizierte Benchmark-Zeile.
Karte - LongCat-2.0
Meituan enthüllte LongCat-2.0 und identifizierte OpenRouters Stealth-Route Owl Alpha als LongCat-gestützt; GitHub/Hugging-Face-Seiten sagen, dass Modellgewichte bald folgen.
Modell
2026-W26
- Studenten-Recherche-Workflow in der Guides-Säule
Pfad, Stack, Playbook und Showcase, um breite Prompts in quellenbasierte Briefings zu verwandeln — der Workflow, den wir neuen Leserinnen zuerst zeigen.
Drop - Factory-Flaggschiff-Dossier veröffentlicht
Agenten-native Auslieferung mit Droid, Spec Mode, Missionen und Enterprise-Kontrollen — Fit 74 mit dokumentierten Autonomie- und Preis-Ausnahmen.
Drop - Cursor-Flaggschiff-Dossier veröffentlicht
VS-Code-kompatibler KI-Editor bei Fit 81 — Tab, Chat, Agent, Privacy Mode, Team-Governance und Diff-Review-Fehlermodi, die wir vor dem Merge dokumentieren.
Drop - Gemini-Flaggschiff-Dossier veröffentlicht
Googles multimodales Assistenten-Ökosystem bei Fit 75 — Suchfundierung, Deep Research, Speicher-Bundles und wo es keine zitierte Suchmaschine ist.
Drop
2026-W25
- NotebookLM-Flaggschiff-Dossier veröffentlicht
Quellengestützter Lese-Workspace bei Fit 78 — Korpus-Chat, Audio-Übersichten, Lernkarten und die leere-Korpus-Falle auf der Karte dokumentiert.
Drop - Perplexity-Flaggschiff-Dossier veröffentlicht
Zitierte Such-Eingangstür bei Fit 72 — Consumer Pro/Max, Deep Research, Sonar-APIs und Fehlermodi für Bibliografien, die du nicht geöffnet hast.
Drop - Cursor
Production-readiness audit pass: editorsNote and pricing/privacy checked dates confirmed. benchmarkReady remains false intentionally until agent-diff safety, usage-burn, and Privacy Mode runs land.
Karte - Factory
Production-readiness audit pass: editorsNote and pricing/privacy checked dates confirmed. benchmarkReady remains false intentionally until Spec Mode, Droid Exec autonomy, and enterprise-control runs land.
Karte - Gemini
Production-readiness audit pass: editorsNote and pricing/privacy checked dates confirmed. No editorial claims changed; bench rows remain explicitly pending until the desk protocol runs (Next action in qualityGate.notes).
Karte - Perplexity
Production-readiness audit pass: editorsNote, pricing/privacy checked dates confirmed; Sonar/Search API token + per-request fee rows aligned with docs.perplexity.ai. benchmarkReady remains false intentionally until desk citation audit and Sonar latency runs land.
Karte - Command Code
Command-Code-Flaggschiff-Karte veröffentlicht — Terminal-Agent mit Taste-Lernen, Open-Model-Deals und vollständiger Preis-/Datenschutz-Recherche.
Karte - Command-Code-Flaggschiff-Dossier veröffentlicht
Terminal-Coding-Agent mit taste-1-Lernen, Open-Model-Deals ab 1 $/Monat und vollständigem Preis- und Datenschutzbild — inklusive .commandcode/.
Drop - OpenCode
OpenCode-Flaggschiff-Karte — OSS-Agent, Go/Zen-Preise, Plan-/Build-Modi und Zen-Datenschutz-Ausnahmen dokumentiert.
Karte - OpenCode-Flaggschiff-Dossier veröffentlicht
Der Open-Source-Coding-Agent bei Fit 95 — Plan-/Build-Modi, Go- und Zen-Preise, 75+ Anbieter und Free-Model-Datenschutz-Ausnahmen dokumentiert.
Drop - Raycast
Raycast als Flaggschiff-Karte für Knowledge work — der macOS-Launcher, den wir auf jedem Forschungs-Mac installieren würden, bevor wir über einen weiteren KI-Tab debattieren.
Karte - Raycast-Flaggschiff-Dossier veröffentlicht
Der macOS-Launcher, den wir installieren, bevor wir über einen weiteren KI-Tab debattieren — Extensions, Quick AI, Zwischenablage und ehrliche Fehlermodi bei Fit 97.
Drop
2026-W24
- GLM-5.2
GLM-5.2 mit 1M-Kontext, 128K Max-Output, zwei Thinking-Modi und Long-Horizon-Coding-Fokus veröffentlicht. MIT-Open-Weights für die folgende Woche angekündigt.
Modell - Kimi K2.7 Code
Kimi K2.7 Code mit 256K-Kontext, verpflichtendem Thinking-Modus, multimodalen Tool-Beispielen und berichteten 30 % weniger Reasoning-Token gegenüber K2.6 veröffentlicht.
Modell - Canva Business
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - ChatGPT
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - ChatPRD
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Claude Fable 5
Claude Fable 5 veröffentlicht — erstes GA Mythos-class-Modell; Vals-Index 75,15 %.
Modell - Connected Papers
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Consensus
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Crossref
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - DeepSeek
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - ElevenLabs
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Elicit
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Exa
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Firecrawl
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Gamma
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Google Scholar
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Grammarly
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Gumloop
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Kagi
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Linear
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Litmaps
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Lovable
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Magic Patterns
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Manus
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Mendeley
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Microsoft Copilot
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Mistral Le Chat
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Mobbin
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - n8n
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Obsidian
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - OpenAlex
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - OpenRouter
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Overleaf
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Paperpile
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - PostHog
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Rayyan
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Replit
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Research Rabbit
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Scholarcy
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Scira AI
Refreshed pricing evidence from scira.ai/about and api.scira.ai/pricing, added machine-readable pricing tiers, and marked benchmark evidence ready for the public dossier while keeping individual desk benchmark rows explicit about what still needs live runs.
Karte - Scite
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Semantic Scholar
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Warp
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Wispr Flow
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - You.com
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Zotero
Converted imported Notion research into a full flagship-ready dossier template with metrics, panes, pricing deck, scenarios, benchmark rows, and comparison slices.
Karte - Förderantrag-Set veröffentlicht
Ein Flagship-Pfad, Stack, Playbook und Walkthrough — von der leeren Aims-Seite zu einreichungsfähigem One-Pager und Budgetbegründung in unter drei Stunden.
Drop
2026-W23
- ChatGPT
ChatGPT ist das/die baseline general assistant card, but it now needs to be framed as a family of surfaces: consumer ChatGPT, Business/Enterprise workspaces, Codex/agent features, and the separate OpenAI API. Its strength is breadth — drafting, file work, multimodal analysis, coding, agents, and extensions — while its risk is that fluent output can make weak sourcing, privacy assumptions, or plan limits invisible.
Karte - Claude
Claude is the careful-drafting and long-context card. The updated evidence reinforces its core lane: nuanced rewriting, document analysis, Projects, Artifacts, Research, Claude Code, Cowork, web search, and organization search/connectors. The editorial risk is quota and surface confusion: consumer Pro/Max, Team, Enterprise, and API/tool pricing are meaningfully different products.
Karte - Connected Papers
Connected Papers bleibt die saubere visuelle Literature-Map-Karte. Besonders nützlich zur Orientierung um ein Seed-Paper, aber die Kernlimitation bleibt methodisch: ein Graph ist kein reproduzierbarer Suchkorpus.
Karte - Consensus
Consensus bleibt die behauptungsförmige akademische Q&A-Karte. Am stärksten, wenn der Nutzer eine Forschungsfrage hat und schnell zitierte Orientierung braucht.
Karte - Crossref
Crossref ist wissenschaftliche Infrastruktur, keine glänzende Discovery-App. Es bleibt die DOI-Metadaten-Verifikationskarte.
Karte - Cursor
Cursor remains the AI-native code editor flagship, but the card now needs to emphasize three separations: editor workflow versus headless agents, individual usage pools versus team/enterprise governance, and Privacy Mode versus ordinary provider transit. Cursor is excellent for repo-local Tab, Chat, Agent, MCP, cloud agents, and Bugbot workflows, but only if diffs are reviewed and tests run.
Karte - DeepSeek
DeepSeek is the low-cost reasoning/API card. This pass sharpens the key facts: DeepSeek’s API is OpenAI- and Anthropic-compatible, current V4-Flash/V4-Pro pricing is extremely low, context length is listed at 1M, and legacy deepseek-chat / deepseek-reasoner names are being deprecated. The low price is real, but the editorial warning is bigger than usual: privacy, residency, institutional policy, and geopolitical availability need explicit review.
Karte - ElevenLabs
ElevenLabs ist die Sprach- und Audio-KI-Karte. Die Produktqualität kann Ergebnisse ausgereift erscheinen lassen.
Karte - Elicit
Elicit ist nicht länger nur ein Paper-Such-Assistent; es positioniert sich als Evidenzsynthese-System für Berichte, systematische Reviews und satzgenaue Zitationen.
Karte - Exa
Exa is one of the cleanest “search as infrastructure” cards in the VerdictPal stack. The current pricing and docs make the value clearer: real-time search, webpage text/highlights, configurable latency, contents retrieval, Answer endpoint, Deep Search, Monitors, and Agent runs. The core risk is economic and privacy-boundary creep when agents fan out across many searches, contents calls, summaries, and enrichment steps.
Karte - Factory
Factory is the agent-native software-delivery card. It is strongest when the task is bigger than autocomplete: repo understanding, spec-planned refactors, terminal/IDE delegation, background agents, SDK use, and CI-style Droid Exec workflows. This pass adds current 2026 Factory/Droid source detail and sharpens the enterprise/security boundary.
Karte - Firecrawl
Firecrawl ist die Web-Extraktions-Infrastrukturkarte. Sollte als Kontext-Extraktion für Agenten und Forschungspipelines beschrieben werden, nicht als generische Suchmaschine.
Karte - Gamma
Gamma ist die KI-Präsentations- und Visual-Document-Erstdraft-Karte. Die Preis- und Datenschutzseiten waren beim Laden nicht erreichbar.
Karte - Gemini
Deepened from official Google pages: pricing across consumer (AI Plus/Pro/Ultra), Workspace add-on, Gemini API paid tier, and Vertex AI; privacy language for consumer vs API vs Workspace/Vertex; capability list refreshed to current Gemini app + API surface.
Karte - Google Scholar
Google Scholar bleibt die akademische Such-Baseline: breit, vertraut, schnell und nützlich zum Finden bekannter Einträge, PDFs, Bibliothekslinks, Zitationsspuren, Alerts und Rechtsprechung.
Karte - Grammarly
Grammarly ist die Überarbeitungsdurchlauf-Karte, nicht die Argument-Schreib-Karte. Dieser Durchlauf aktualisiert die Karte für Grammarly Pro und Superhuman-Suite.
Karte - Gumloop
Gumloop ist die No-Code-KI-Workflow- und Agent-Automation-Karte. Basiert auf primären Preis- und Datenschutzseiten plus offiziellen Produkt-/Suchergebnissen.
Karte - Kagi
Kagi ist die Paid-Search-Karte: eine nutzerfinanzierte, werbefreie Suchmaschine mit ernsthafter Datenschutz-Technik und nützlicher Assistant-Funktionalität.
Karte - Litmaps
Litmaps ist die Citation-Map-plus-Monitoring-Karte. Wertvoll, wenn Nutzer von Seed-Papern ausgehen und vernetzte Arbeiten entdecken, visualisieren und überwachen wollen.
Karte - Mendeley
Mendeley ist der Elsevier-gestützte Cloud-Sync-Referenzmanager mit zunehmend KI-geprägten Bibliotheksfunktionen.
Karte - Microsoft Copilot
Microsoft Copilot ist a family-name card, not a single-product card. This pass uses Microsoft Learn pages for Microsoft 365 Copilot and Copilot Chat data protection plus Microsoft pricing search results. The key editorial job is to separate consumer Copilot, Microsoft 365 Copilot Chat, paid Microsoft 365 Copilot, and Azure/OpenAI developer paths.
Karte - Mistral Le Chat
Mistral Le Chat ist is the EU-headquartered assistant comparison card. This pass strengthens the privacy and API boundary: Le Chat inputs/outputs are retained until account/conversation deletion, API inputs/outputs are generally kept for 30 rolling days for abuse monitoring unless zero data retention is activated, and Agents/Fine-tuning have separate retention rules. That makes Mistral useful to compare against US-default assistants, but not automatically compliant.
Karte - n8n
n8n ist die Research-Ops-Automatisierungskarte. Am stärksten, wenn ein Team sichtbaren Workflow-Kleber über APIs, Webhooks, Datenbanken, KI-Schritte, Notion und Alerts benötigt.
Karte - NotebookLM
NotebookLM remains the source-grounded workspace flagship. The strongest VerdictPal framing is not “AI search,” but “corpus-first reading workspace”: upload or discover sources, ask grounded questions, generate study artifacts, and verify claims against citations. This pass adds current Enterprise, Workspace, and Google AI plan distinctions so the privacy story is clearer.
Karte - Obsidian
Obsidian ist die Local-First-Wissensdatenbank-Karte. Das wichtige Update: Obsidian ist jetzt sowohl für die Arbeit als auch privat kostenlos; Sync und Publish bleiben optional.
Karte - OpenAlex
OpenAlex ist die offene wissenschaftliche Graph-Karte. Sie sitzt zwischen Crossref und Semantic Scholar: breitere Discovery als DOI-Registry, reproduzierbarer als geschlossene Suchprodukte.
Karte - OpenRouter
OpenRouter ist die Model-Router-Karte. Nützlich, wenn ein Entwickler eine OpenAI-kompatible API-Oberfläche für viele LLMs, Fallback-Routing, Modellvergleiche und BYOK möchte.
Karte - Overleaf
Overleaf ist die kollaborative LaTeX-Karte. Dieser Durchlauf aktualisiert die Karte mit aktuellen Plan-Nachweisen: KI-Kontingent, Mitarbeiterlimits und 24x-Kompilierzeit.
Karte - Paperpile
Paperpile ist der Convenience-First-Referenzmanager für Google Docs und browser-natives Schreiben. Exzellente Google-Docs- und PDF-Workflow-Abdeckung, aber Tradeoffs bei bibliographischer Tiefe.
Karte - Perplexity
Perplexity remains the cited-answer benchmark for consumer AI search, but the card now needs to treat Perplexity as two related products: the consumer answer engine and the API platform. The consumer product is useful for fast source maps and Deep Research drafts; the API platform is a separate developer surface with Sonar, Search, Agent, and Embeddings APIs. The hard editorial rule stays the same: citations are leads, not bibliography-ready evidence.
Karte - PostHog
PostHog ist die Produktanalyse- und Experimentier-Stack-Karte. Sie gehört in VerdictPal, weil Produktanalyse, Session Replay, Feature Flags und Experimente sich mit Forschungs- und Produkt-Ops überschneiden.
Karte - Rayyan
Rayyan ist ein fokussierter Systematic-Review-Screening-Workspace mit stärkeren Plan-Mechaniken.
Karte - Replit
Replit ist die browserbasierte KI-Build-and-Deploy-Karte. Nützlich, wenn ein Nutzer einen Ort für Ideenfindung, Agent-Building, Datenbank, Hosting, Zusammenarbeit und Publishing möchte.
Karte - Research Rabbit
Research Rabbit ist the visual exploration and “follow the trail” card. It belongs beside Connected Papers and Litmaps, but its emphasis is adaptive exploration, collections, and seeing how papers/authors/concepts connect. This pass adds current feature language, data-source hints, and a clearer privacy/method boundary.
Karte - Scholarcy
Scholarcy ist die Paper-Überflug- und strukturierte-Zusammenfassungs-Karte. Kann Zeit für Triage, Accessibility und Organisation sparen, muss aber als Lesehilfe geprüft werden.
Karte - Scira AI
Scira remains the VerdictPal flagship template because it combines a cited search UI, open-source AGPL codebase, hosted Free/Pro/Max plans, an API platform, MCP surface, and a deep dossier structure. This pass replaces stale/malformed page content with a cleaner benchmarkable dossier while preserving the core stance: strong public recommendation, but benchmark-ready still false until the pending desk rows are actually run.
Karte - Scite
Scite ist die Zitationskontext-Karte. Wertvoll, weil sie eine bessere Frage stellt: Haben spätere Paper die Behauptung gestützt, kontrastiert oder nur erwähnt?
Karte - Semantic Scholar
Semantic Scholar bleibt eine Flagship-akademische-Suchschicht, weil sie eine kostenlose wissenschaftliche Such-UI mit TLDRs, Autorenseiten, Alerts und Entwickler-API kombiniert.
Karte - Warp
Warp ist die agentische Terminal/Workspace-Karte. Überzeugend, weil sie Coding-Agenten ins Terminal bringt.
Karte - Zotero
Zotero bleibt der stärkste Local-First-Standard für studentisches und akademisches Referenzmanagement.
Karte - Canva Business
Aus der Notion-Karten-Pipeline importiert mit Scira-gate status corrected to solid.
Karte - ChatPRD
Notion-State-of-Art-Recherche importiert und redaktionelle Eignung gesenkt, um schwache Desk-Nützlichkeit plus Teams-gesicherte Linear-Integration widerzuspiegeln.
Karte - Magic Patterns
Aus der Notion-Karten-Pipeline importiert mit Scira-gate status corrected to solid.
Karte - Dark Mode für das Evidence Instrument überarbeitet
Dunkle Oberflächen nutzen jetzt gedämpftes Amber, lesbare Ränder und komponentenspezifische Overrides statt Neon-Karten und invertierter Cream-Panels.
Drop - Trusted-Source-Briefing-Workflow veröffentlicht
Die Guides-Säule hat jetzt einen echten Recherchepfad, Tool-Stack, Playbook und Showcase, um breite Prompts in quellenbasierte Briefings zu verwandeln.
Drop - Claude Mythos Preview
Claude-Mythos-Research-Preview veröffentlicht — Spitze bei Vals (73,42 %) und SWE-bench Verified (93,9 %, Best-of-3).
Modell
2026-W22
2026-W21
- Gemini 3.5 Flash
Gemini 3.5 Flash veröffentlicht — MCP-Atlas-Spitze (83,6 %), 1M-Kontext, günstigster Frontier-Preis.
Modell
2026-W18
2026-W17
- DeepSeek V4 Pro Max
DeepSeek V4 Pro Max mit Open Weights und Frontier-Coding-Scores veröffentlicht.
Modell - Kimi K2.6
Kimi K2.6 veröffentlicht — Vals #5 bei 55,55 %, GPQA Diamond 90,5 % (Open-Source-Spitze).
Modell
2026-W16
- Claude Opus 4.7
Claude Opus 4.7 veröffentlicht — 13 % Coding-Steigerung, 3× Produktionsaufgaben, 3,75 MP Vision.
Modell
2026-W15
- GLM-5.1
GLM-5.1 mit 200K-Kontext, 128K Max-Output, Long-Horizon-Coding-Fokus und MIT-lizenzierten Open Weights veröffentlicht.
Modell
2026-W14
2026-W12
2026-W11
2026-W10
2026-W08
- Claude Sonnet 4.6
Claude Sonnet 4.6 veröffentlicht — Vals #3 bei 60,30 %, 1M-Kontext, $3 / $15.
Modell
2026-W07
2026-W06
- Claude Opus 4.6
Claude Opus 4.6 veröffentlicht — erstes Opus mit 1M-Kontext, 128K-Output, Agent-Teams.
Modell
2025-W50
2025-W48
- Claude Opus 4.5
Claude Opus 4.5 veröffentlicht — bestes Coding- und Agent-Modell bei Launch, $5/$25.
Modell