Why this benchmark is useful
For source-heavy students and researchers this is the live science-agent discriminator: real lab-style workflows, still far from ceiling (best published 0.1 row is 30%).
TB-Science
Whether an agent can finish real scientific research workflows in a terminal: 70 expert-curated tasks across life, physical, Earth, mathematical, and engineering sciences, graded on concrete artifacts (analyses, simulations, proofs, code, data products).
For source-heavy students and researchers this is the live science-agent discriminator: real lab-style workflows, still far from ceiling (best published 0.1 row is 30%).
Read resolution rate with the harness named on the row. Claude Opus 5 + Claude Code leads the 0.1 table at 30.0%; GPT-5.6 Sol + Codex is 22.4%. GLM-5.3 is the strongest open-weight row at 8.1%. Do not compare these to Terminal-Bench 2.1 coding headlines.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Opus 5AnthropicOfficial 0.1 table: Claude Code harness, three trials × 70 tasks. Do not mix with Terminal-Bench 2.1. | 2026-07-24 | 30% | Terminal-Bench-Science 0.1 announcement2026-08-27 | Source-checked |
| 2 | GPT-5.6 SolOpenAIOfficial 0.1 table: Codex harness. Matches Fable 5 resolution at less than a third of the announced eval cost. | 2026-07-09 | 22.4% | Terminal-Bench-Science 0.1 announcement2026-08-27 | Source-checked |
| 3 | Claude Fable 5Anthropic | 2026-06-09 | 21.4% | Terminal-Bench-Science 0.1 announcement2026-08-27 | Source-checked |
| 4 | Claude Opus 4.8Anthropic | 2026-05-28 | 10.5% | Terminal-Bench-Science 0.1 announcement2026-08-27 | Source-checked |
| 5 | GPT-5.6 TerraOpenAI | 2026-07-09 | 8.6% | Terminal-Bench-Science 0.1 announcement2026-08-27 | Source-checked |
| 6 | GLM-5.3ZhipuStrongest open-weight row on the 0.1 announcement table. | 2026-08-14 | 8.1% | Terminal-Bench-Science 0.1 announcement2026-08-27 | Source-checked |
| 7 | Kimi K3Moonshot | 2026-07-16 | 7.1% | Terminal-Bench-Science 0.1 announcement2026-08-27 | Source-checked |
| 7 | Grok 4.6xAI | 2026-08-12 | 7.1% | Terminal-Bench-Science 0.1 announcement2026-08-27 | Source-checked |
| 9 | GPT-5.6 LunaOpenAI | 2026-07-09 | 3.3% | Terminal-Bench-Science 0.1 announcement2026-08-27 | Source-checked |
Scores from the Terminal-Bench-Science 0.1 announcement (27 Aug 2026): three trials per task, 70 tasks. Work on 0.2 is open with a 5 Oct 2026 PR deadline. Official Terminal-Bench 4.0 is a different, software-engineering suite — not this page.
Human expert time is research-workflow scale. Published 0.1 resolution rates cluster at or below 30%.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
Whether an agent can finish real scientific research workflows in a terminal: 70 expert-curated tasks across life, physical, Earth, mathematical, and engineering sciences, graded on concrete artifacts (analyses, simulations, proofs, code, data products).
A strong Terminal-Bench-Science result says nothing about:
Resolution rate across 70 tasks (pass if the hidden checker accepts the artifact). VerdictPal rows use the 27 Aug 2026 announcement table, one row per model at its published harness. Task format: 70 containerized Harbor tasks. Agents get a terminal, files, and tools; graders check the produced artifact with task-specific tests. Three independent trials per task in the 0.1 announcement table.
Terminal-Bench-Science is currently marked Active in the atlas. Ceiling context: Human expert time is research-workflow scale. Published 0.1 resolution rates cluster at or below 30%.
Contamination risk for Terminal-Bench-Science is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.