Benchmarks / Agentic

Terminal-Bench-Science

TB-Science

Whether an agent can finish real scientific research workflows in a terminal: 70 expert-curated tasks across life, physical, Earth, mathematical, and engineering sciences, graded on concrete artifacts (analyses, simulations, proofs, code, data products).

What this does not measure
  • The model alone — every published 0.1 row names an agent harness (Claude Code, Codex, Grok Build). Harness choice moves the score.
  • Hypothesis generation or paper writing — tasks are executable workflows with hidden checkers, not open-ended discovery.
  • Terminal-Bench 2.1 software-engineering skill — TB-Science is a separate, harder science suite. Do not mix 2.1 coding pass rates with 0.1 science resolution.
Analysis

Why this benchmark is useful

For source-heavy students and researchers this is the live science-agent discriminator: real lab-style workflows, still far from ceiling (best published 0.1 row is 30%).

Scope

Coverage map

Task family
Agentic
Format
70 containerized Harbor tasks. Agents get a terminal, files, and tools; graders check the produced artifact with task-specific tests. Three independent trials per task in the 0.1 announcement table.
Scoring
Resolution rate across 70 tasks (pass if the hidden checker accepts the artifact). VerdictPal rows use the 27 Aug 2026 announcement table, one row per model at its published harness.
Maintainer
Stanford / Laude Institute / Terminal-Bench
Reading guide

How to read the scores

Read resolution rate with the harness named on the row. Claude Opus 5 + Claude Code leads the 0.1 table at 30.0%; GPT-5.6 Sol + Codex is 22.4%. GLM-5.3 is the strongest open-weight row at 8.1%. Do not compare these to Terminal-Bench 2.1 coding headlines.

Blind spots

What it does not cover

  • The model alone — every published 0.1 row names an agent harness (Claude Code, Codex, Grok Build). Harness choice moves the score.
  • Hypothesis generation or paper writing — tasks are executable workflows with hidden checkers, not open-ended discovery.
  • Terminal-Bench 2.1 software-engineering skill — TB-Science is a separate, harder science suite. Do not mix 2.1 coding pass rates with 0.1 science resolution.
Scores

Evidence ledger

9 rows
99 rows
30%best score
9source-checked
1sources
2026-08-27source date
Terminal-Bench9

official-page · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Opus 530% · Terminal-Bench-Science 0.1 announcement
  2. GPT-5.6 Sol22.4% · Terminal-Bench-Science 0.1 announcement
  3. Claude Fable 521.4% · Terminal-Bench-Science 0.1 announcement
  4. Claude Opus 4.810.5% · Terminal-Bench-Science 0.1 announcement
  5. GPT-5.6 Terra8.6% · Terminal-Bench-Science 0.1 announcement

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Opus 5AnthropicOfficial 0.1 table: Claude Code harness, three trials × 70 tasks. Do not mix with Terminal-Bench 2.1.2026-07-2430%
Terminal-Bench-Science 0.1 announcement2026-08-27
Source-checked
2GPT-5.6 SolOpenAIOfficial 0.1 table: Codex harness. Matches Fable 5 resolution at less than a third of the announced eval cost.2026-07-0922.4%
Terminal-Bench-Science 0.1 announcement2026-08-27
Source-checked
3Claude Fable 5Anthropic2026-06-0921.4%
Terminal-Bench-Science 0.1 announcement2026-08-27
Source-checked
4Claude Opus 4.8Anthropic2026-05-2810.5%
Terminal-Bench-Science 0.1 announcement2026-08-27
Source-checked
5GPT-5.6 TerraOpenAI2026-07-098.6%
Terminal-Bench-Science 0.1 announcement2026-08-27
Source-checked
6GLM-5.3ZhipuStrongest open-weight row on the 0.1 announcement table.2026-08-148.1%
Terminal-Bench-Science 0.1 announcement2026-08-27
Source-checked
7Kimi K3Moonshot2026-07-167.1%
Terminal-Bench-Science 0.1 announcement2026-08-27
Source-checked
7Grok 4.6xAI2026-08-127.1%
Terminal-Bench-Science 0.1 announcement2026-08-27
Source-checked
9GPT-5.6 LunaOpenAI2026-07-093.3%
Terminal-Bench-Science 0.1 announcement2026-08-27
Source-checked
Method

What it covers

Scores from the Terminal-Bench-Science 0.1 announcement (27 Aug 2026): three trials per task, 70 tasks. Work on 0.2 is open with a 5 Oct 2026 PR deadline. Official Terminal-Bench 4.0 is a different, software-engineering suite — not this page.

Score ceiling

Where it breaks down

Human expert time is research-workflow scale. Published 0.1 resolution rates cluster at or below 30%.

Tools that report it

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does Terminal-Bench-Science measure?

Whether an agent can finish real scientific research workflows in a terminal: 70 expert-curated tasks across life, physical, Earth, mathematical, and engineering sciences, graded on concrete artifacts (analyses, simulations, proofs, code, data products).

What does a high Terminal-Bench-Science score not prove?

A strong Terminal-Bench-Science result says nothing about:

  • The model alone — every published 0.1 row names an agent harness (Claude Code, Codex, Grok Build). Harness choice moves the score.
  • Hypothesis generation or paper writing — tasks are executable workflows with hidden checkers, not open-ended discovery.
  • Terminal-Bench 2.1 software-engineering skill — TB-Science is a separate, harder science suite. Do not mix 2.1 coding pass rates with 0.1 science resolution.

How is Terminal-Bench-Science scored?

Resolution rate across 70 tasks (pass if the hidden checker accepts the artifact). VerdictPal rows use the 27 Aug 2026 announcement table, one row per model at its published harness. Task format: 70 containerized Harbor tasks. Agents get a terminal, files, and tools; graders check the produced artifact with task-specific tests. Three independent trials per task in the 0.1 announcement table.

Is Terminal-Bench-Science saturated?

Terminal-Bench-Science is currently marked Active in the atlas. Ceiling context: Human expert time is research-workflow scale. Published 0.1 resolution rates cluster at or below 30%.

Can Terminal-Bench-Science results be contaminated by training data?

Contamination risk for Terminal-Bench-Science is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.