46.8%
AA-Omniscience Accuracy
Models / OpenAI
GPT-5.6 · Released 2026-07-09
OpenAI's generally available balanced GPT-5.6 model, positioned below Sol for efficient everyday work at $2.50 / $15 per 1M tokens.
Balanced GPT-5.6 tier for daily coding, research, and agent workflows when Sol's higher price is unnecessary.
Family profile
Best published score in each covered benchmark family.
6 tested benchmark families
Where GPT-5.6 Terra places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.
The three highest-scoring models with pages in each capability family. Where GPT-5.6 Terra shows up, it's highlighted.
Benchmark rows added to the public ledger for GPT-5.6 Terra in the last 120 days. Older rows live in the full table below.
| Benchmark | Family | Score | Source | Days ago |
|---|---|---|---|---|
| CritPt | Reasoning | 30% | Artificial Analysis — CritPt | 5 |
| AA-Omniscience | Knowledge | 46.8% | Artificial Analysis — AA-Omniscience | 5 |
| MMMU-Pro | Multimodal | 80.7% | Artificial Analysis — MMMU-Pro | 5 |
| AA-LCR | Long context | 79.67% | Artificial Analysis | 5 |
| AA output speed | Performance | 108 t/s | Artificial Analysis | 5 |
| AA time to first token | Performance | 94.25s | Artificial Analysis | 5 |
| Artificial Analysis Intelligence Index | Reasoning | 56.6% | Artificial Analysis | 5 |
| GPQA Diamond | Reasoning | 92.5% | Artificial Analysis | 5 |
| Humanity's Last Exam | Reasoning | 42.9% | Artificial Analysis | 5 |
| IFBench | Reasoning | 71.22% | Artificial Analysis | 5 |
| SciCode | Coding | 53.9% | Artificial Analysis | 5 |
| τ³-Bench Banking | Agentic | 40.21% | Artificial Analysis | 5 |
| Terminal-Bench | Agentic | 88.01% | Artificial Analysis | 5 |
| Terminal-Bench-Science | Agentic | 8.6% | Terminal-Bench-Science 0.1 announcement | 10 |
| Vending-Bench 2 | Agentic | 7343.21 | Andon Labs Vending-Bench 2 | 33 |
A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.
Hover any dot for name, score, and input price. Keyboard: tab through the top twelve, or use the ranking below.
Every catalog benchmark for GPT-5.6 Terra. Scores link to the original source; gaps mean no public row exists yet.
15 of 37 catalog benchmarks have a sourced row for GPT-5.6 Terra.
46.8%
AA-Omniscience Accuracy
56.6%
AA Intelligence Index v4.1
30%
CritPt composite (AA run)
92.5%
GPQA Diamond (AA run)
42.9%
HLE (AA run)
71.22%
IFBench (AA run)
53.9%
SciCode (AA run)
88.01%
Terminal-Bench v2.1 (AA run)
8.6%
TB-Science 0.1 resolution (Codex)
7343.21
Final bank balance (USD)
40.21%
τ³-Bench Banking (AA run)
80.7%
MMMU-Pro (AA run)
79.67%
AA-LCR (AA run)
108 t/s
Median output tokens/s (1k prompt, default provider)
94.25s
Median time to first token (1k prompt, default provider)