Benchmarks / Tool use

MCP Atlas

How reliably a model selects, sequences, and recovers when invoking tools exposed via the Model Context Protocol — including schema adherence, error handling, and chain-of-tool reasoning.

What this does not measure
  • Harness quality — different MCP server implementations can artificially boost or suppress scores.
  • Cost — a model that hammers a tool until it succeeds is scored the same as a model that calls it cleanly once.
  • Real production reliability — these are clean lab tasks, not the messy recovery paths of a long-lived agent.
Analysis

Why this benchmark is useful

When you care whether a model can select, sequence, and recover tool calls over MCP — schema adherence and chain-of-tool reasoning, not chat polish.

Scope

Coverage map

Task family
Tool use
Format
Scored tool-call sequences against an MCP server; the suite measures both call correctness and recovery from intermediate errors.
Scoring
Percentage of tasks completed end-to-end; see the source repo for harness and scoring script details.
Maintainer
MCP-Atlas community
Reading guide

How to read the scores

End-to-end task completion percentage against an MCP server. Harness and server implementation can swing scores; cost of retries is not scored.

Blind spots

What it does not cover

  • Harness quality — different MCP server implementations can artificially boost or suppress scores.
  • Cost — a model that hammers a tool until it succeeds is scored the same as a model that calls it cleanly once.
  • Real production reliability — these are clean lab tasks, not the messy recovery paths of a long-lived agent.
Scores

Evidence ledger

10 rows
1010 rows
83.6%best score
1source-checked
1sources
2026-06-09to 2026-05-15
Papers / model cards10

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Opus 4.879.2% · MCP Atlas leaderboard
  2. GPT 5.577.8% · MCP Atlas leaderboard
  3. Claude Sonnet 4.675.4% · MCP Atlas leaderboard
  4. Gemini 3.1 Pro74.1% · MCP Atlas leaderboard
  5. DeepSeek V4 Pro Max72.6% · MCP Atlas leaderboard

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Gemini 3.5 FlashGoogle2026-05-1983.6%
MCP Atlas leaderboard2026-05-15
Source-checked
2Claude Opus 4.8Anthropic2026-05-2879.2%
MCP Atlas leaderboard2026-06-09
Needs audit
3GPT 5.5OpenAI2026-04-2377.8%
MCP Atlas leaderboard2026-06-09
Needs audit
4Claude Sonnet 4.6Anthropic2026-02-1775.4%
MCP Atlas leaderboard2026-06-09
Needs audit
5Gemini 3.1 ProGoogle2026-02-1974.1%
MCP Atlas leaderboard2026-06-09
Needs audit
6DeepSeek V4 Pro MaxDeepSeek2026-04-2472.6%
MCP Atlas leaderboard2026-06-09
Needs audit
7Qwen 3.7 MaxAlibaba2026-05-2071.8%
MCP Atlas leaderboard2026-06-09
Needs audit
8Mistral Large 3Mistral2025-12-0270.4%
MCP Atlas leaderboard2026-06-09
Needs audit
9Kimi K2.6Moonshot2026-04-2069.7%
MCP Atlas leaderboard2026-06-09
Needs audit
10MiniMax M3MiniMax2026-05-3168.9%
MCP Atlas leaderboard2026-06-09
Needs audit
Method

What it covers

This benchmark sits in the tool use family. It uses Scored tool-call sequences against an MCP server; the suite measures both call correctness and recovery from intermediate errors. The score should travel with its task format, scoring method, source date, and benchmark version.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Tools that report it

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does MCP Atlas measure?

How reliably a model selects, sequences, and recovers when invoking tools exposed via the Model Context Protocol — including schema adherence, error handling, and chain-of-tool reasoning.

What does a high MCP Atlas score not prove?

A strong MCP Atlas result says nothing about:

  • Harness quality — different MCP server implementations can artificially boost or suppress scores.
  • Cost — a model that hammers a tool until it succeeds is scored the same as a model that calls it cleanly once.
  • Real production reliability — these are clean lab tasks, not the messy recovery paths of a long-lived agent.

How is MCP Atlas scored?

Percentage of tasks completed end-to-end; see the source repo for harness and scoring script details. Task format: Scored tool-call sequences against an MCP server; the suite measures both call correctness and recovery from intermediate errors.

Is MCP Atlas saturated?

MCP Atlas is currently marked Active in the atlas.

Can MCP Atlas results be contaminated by training data?

Contamination risk for MCP Atlas is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.