58.9%
AA Intelligence Index v4.1
Modelle / OpenAI
GPT-5.6 · Release 2026-07-09
OpenAIs allgemein verfügbares GPT-5.6-Flagship für Coding, Wissenschaft, Cybersecurity, Computer-Nutzung und langlaufende agentische Arbeit, zum Preis von $5 / $30 pro 1 Mio. Token.
GPT-5.5-Nachfolger für Frontier-Coding, Research, Computer-Nutzung und Multi-Agent-Workflows, jetzt mit öffentlichen Benchmark-Zeilen von OpenAI, Vals und Artificial Analysis.
78average across 4 tested families
Wo GPT-5.6 Sol auf jeder öffentlichen Benchmark-Quelle mit Zeile steht. Rang zählt jedes Modell mit neuester Zeile im selben Test — kein universeller Qualitätsscore.
| Benchmark | Family | Rang | Score | Quelle | Datum |
|---|---|---|---|---|---|
| Terminal-Bench | Agentic | #1/ 69 | 88 | Artificial Analysis | 2026-07-21 |
| Artificial Analysis Intelligence Index | Reasoning | #2/ 67 | 59 | Artificial Analysis | 2026-07-21 |
| Humanity's Last Exam | Reasoning | #2/ 65 | 47 | Artificial Analysis | 2026-07-21 |
| GPQA Diamond | Reasoning | #3/ 69 | 94 | Artificial Analysis | 2026-07-21 |
| Vals Index | Agentic | #4/ 20 | 73 | Vals AI — Vals Index | 2026-07-16 |
| AA time to first token | Performance | #6/ 69 | Artificial Analysis | 2026-07-21 | |
| AA-LCR | Long context | #8/ 64 | 74 | Artificial Analysis | 2026-07-21 |
| SciCode | Coding | #8/ 65 | 56 | Artificial Analysis | 2026-07-21 |
| τ³-Bench Banking | Agentic | #21/ 64 | 33 | Artificial Analysis | 2026-07-21 |
| IFBench | Reasoning | #27/ 57 | 73 | Artificial Analysis | 2026-07-21 |
| AA output speed | Performance | #42/ 69 | Artificial Analysis | 2026-07-21 |
Die drei höchstscorierenden kartierten Modelle pro Capability-Family. Wo GPT-5.6 Sol auftaucht, ist es hervorgehoben.
Benchmark-Zeilen im öffentlichen Ledger für GPT-5.6 Sol in den letzten 120 Tagen. Ältere Zeilen stehen in der vollen Tabelle unten.
| Benchmark | Family | Score | Quelle | Tage her |
|---|---|---|---|---|
| AA-LCR | Long context | 73.67% | Artificial Analysis | 2 |
| AA output speed | Performance | 67 t/s | Artificial Analysis | 2 |
| AA time to first token | Performance | 76.91s | Artificial Analysis | 2 |
| Artificial Analysis Intelligence Index | Reasoning | 58.9% | Artificial Analysis | 2 |
| GPQA Diamond | Reasoning | 94.1% | Artificial Analysis | 2 |
| Humanity's Last Exam | Reasoning | 47.2% | Artificial Analysis | 2 |
| IFBench | Reasoning | 72.65% | Artificial Analysis | 2 |
| SciCode | Coding | 56.1% | Artificial Analysis | 2 |
| τ³-Bench Banking | Agentic | 32.99% | Artificial Analysis | 2 |
| Terminal-Bench | Agentic | 88.01% | Artificial Analysis | 2 |
| Vals Index | Agentic | 73.1% | Vals AI — Vals Index | 7 |
Ein hoher Composite, der eine schwache Family versteckt, ist eine Falle. Diese Balken zeigen Families ohne öffentlichen Test — und wo das Modell führt.
Hover or tab any dot for name, score, and input price.
Jeder Katalog-Benchmark für GPT-5.6 Sol. Scores verlinken zur Originalquelle; Lücken heißen: noch keine öffentliche Zeile.
11 of 32 catalog benchmarks have a sourced row for GPT-5.6 Sol.
58.9%
AA Intelligence Index v4.1
94.1%
GPQA Diamond (AA run)
47.2%
HLE (AA run)
72.65%
IFBench (AA run)
56.1%
SciCode (AA run)
88.01%
Terminal-Bench v2.1 (AA run)
73.1%
Vals Index
32.99%
τ³-Bench Banking (AA run)
73.67%
AA-LCR (AA run)
67 t/s
Median output tokens/s (1k prompt, default provider)
76.91s
Median time to first token (1k prompt, default provider)