Benchmarks / Tool use

MCP Atlas

How reliably a model selects, sequences, and recovers when invoking tools exposed via the Model Context Protocol — including schema adherence, error handling, and chain-of-tool reasoning.

Tool useSolidActiveLow contamination riskSince 2026
What this does not measure
  • Harness quality — different MCP server implementations can artificially boost or suppress scores.
  • Cost — a model that hammers a tool until it succeeds is scored the same as a model that calls it cleanly once.
  • Real production reliability — these are clean lab tasks, not the messy recovery paths of a long-lived agent.
Analysis

Why this benchmark is useful

Editorial brief pendingWe publish the methodology and ledger first; benchmark-specific analysis ships after desk review.

Scope

Coverage map

Task family
Tool use
Format
Scored tool-call sequences against an MCP server; the suite measures both call correctness and recovery from intermediate errors.
Scoring
Percentage of tasks completed end-to-end; see the source repo for harness and scoring script details.
Maintainer
MCP-Atlas community
Reading guide

How to read the scores

Reading guide pending. Use task format, scoring method, and source dates in the ledger until the desk brief ships.

Blind spots

What it does not cover

  • Harness quality — different MCP server implementations can artificially boost or suppress scores.
  • Cost — a model that hammers a tool until it succeeds is scored the same as a model that calls it cleanly once.
  • Real production reliability — these are clean lab tasks, not the messy recovery paths of a long-lived agent.
Scores

Evidence ledger

10 rows
1010 rows
83.6%best score
1source-checked
1sources
2026-06-09to 2026-05-15
Papers / model cards10

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Opus 4.879.2% · MCP Atlas leaderboard
  2. GPT 5.577.8% · MCP Atlas leaderboard
  3. Claude Sonnet 4.675.4% · MCP Atlas leaderboard
  4. Gemini 3.1 Pro74.1% · MCP Atlas leaderboard
  5. DeepSeek V4 Pro Max72.6% · MCP Atlas leaderboard

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Gemini 3.5 FlashGoogle2026-05-1983.6%
MCP Atlas leaderboard2026-05-15
Source-checked
2Claude Opus 4.8Anthropic2026-05-2879.2%
MCP Atlas leaderboard2026-06-09
Needs audit
3GPT 5.5OpenAI2026-04-2377.8%
MCP Atlas leaderboard2026-06-09
Needs audit
4Claude Sonnet 4.6Anthropic2026-02-1775.4%
MCP Atlas leaderboard2026-06-09
Needs audit
5Gemini 3.1 ProGoogle2026-02-1974.1%
MCP Atlas leaderboard2026-06-09
Needs audit
6DeepSeek V4 Pro MaxDeepSeek2026-04-2472.6%
MCP Atlas leaderboard2026-06-09
Needs audit
7Qwen 3.7 MaxAlibaba2026-05-2071.8%
MCP Atlas leaderboard2026-06-09
Needs audit
8Mistral Large 3Mistral2025-12-0270.4%
MCP Atlas leaderboard2026-06-09
Needs audit
9Kimi K2.6Moonshot2026-04-2069.7%
MCP Atlas leaderboard2026-06-09
Needs audit
10MiniMax M3MiniMax2026-05-3168.9%
MCP Atlas leaderboard2026-06-09
Needs audit
Method

What it covers

Data-quality note pending. Every ledger row still carries source URL, source date, and ingest timestamp.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Tools that report it

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.