Benchmarks / Tool use

Berkeley Function Calling Leaderboard

BFCL

Whether a model can choose and format function calls correctly across single-turn, multi-turn, parallel, and live API-style tool-use tasks.

Tool useSolidActiveMedium contamination riskSince 2024
What this does not measure
  • Whether the downstream tool result is useful — BFCL focuses on choosing and formatting calls, not on the whole product workflow after the call returns.
  • Security of tool execution — permission boundaries, prompt-injection resistance, and audit logs still need product-specific tests.
  • Long-horizon agent reliability — passing function-call tasks does not prove the system can recover over dozens of steps.
Analysis

Why this benchmark is useful

Tool calling is the plumbing behind agents, automations, and research assistants. BFCL gives that plumbing a public test instead of treating every function-call demo as evidence.

Scope

Coverage map

Task family
Tool use
Format
Function/tool-calling tasks with schema-constrained expected outputs across simple, multi-turn, parallel, and live-style settings.
Scoring
Accuracy-style score by BFCL category and version. The exact category matters more than the aggregate.
Maintainer
Berkeley Gorilla
Reading guide

How to read the scores

Compare models within the same BFCL version and category. A model can be strong at simple calls and weak at parallel or multi-turn calls.

Blind spots

What it does not cover

  • Whether the downstream tool result is useful — BFCL focuses on choosing and formatting calls, not on the whole product workflow after the call returns.
  • Security of tool execution — permission boundaries, prompt-injection resistance, and audit logs still need product-specific tests.
  • Long-horizon agent reliability — passing function-call tasks does not prove the system can recover over dozens of steps.
Scores

Evidence ledger

20 rows
2020 rows
88.7%best score
0source-checked
1sources
2026-04-15source date
Berkeley BFCL20

official-page · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Sonnet 4.688.7% · Berkeley Function Calling Leaderboard V4
  2. GPT 5.587.4% · Berkeley Function Calling Leaderboard V4
  3. Claude Opus 4.886.9% · Berkeley Function Calling Leaderboard V4
  4. Kimi K2.685.2% · Berkeley Function Calling Leaderboard V4
  5. GPT 5.483.8% · Berkeley Function Calling Leaderboard V4

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Sonnet 4.6AnthropicV4 aggregate score; drill into multi-turn and parallel sub-scores before choosing.2026-02-1788.7%
Berkeley Function Calling Leaderboard V42026-04-15
Needs audit
2GPT 5.5OpenAI2026-04-2387.4%
Berkeley Function Calling Leaderboard V42026-04-15
Needs audit
3Claude Opus 4.8Anthropic2026-05-2886.9%
Berkeley Function Calling Leaderboard V42026-04-15
Needs audit
4Kimi K2.6Moonshot2026-04-2085.2%
Berkeley Function Calling Leaderboard V42026-04-15
Needs audit
5GPT 5.4OpenAI2026-03-0583.8%
Berkeley Function Calling Leaderboard V42026-04-15
Needs audit
6Gemini 3.5 FlashGoogle2026-05-1982.3%
Berkeley Function Calling Leaderboard V42026-04-15
Needs audit
7Gemini 3.1 ProGoogle2026-02-1981.6%
Berkeley Function Calling Leaderboard V42026-04-15
Needs audit
8Mistral Large 3Mistral2025-12-0280.1%
Berkeley Function Calling Leaderboard V42026-04-15
Needs audit
9Gemini 3 Flash PreviewGoogle2025-12-1779.4%
Berkeley Function Calling Leaderboard V42026-04-15
Needs audit
10Qwen 3.7 MaxAlibaba2026-05-2078.4%
Berkeley Function Calling Leaderboard V42026-04-15
Needs audit
11o3OpenAIStrong in simple single-turn; drops in multi-turn complexity.2025-04-1678.1%
Berkeley Function Calling Leaderboard V42026-04-15
Needs audit
12Grok 4.3xAI2026-04-3077.2%
Berkeley Function Calling Leaderboard V42026-04-15
Needs audit
13DeepSeek V4 Flash MaxDeepSeek2026-04-2476.8%
Berkeley Function Calling Leaderboard V42026-04-15
Needs audit
14MiniMax M3MiniMax2026-05-3176.2%
Berkeley Function Calling Leaderboard V42026-04-15
Needs audit
15Llama 4 MaverickMeta2026-04-0575.9%
Berkeley Function Calling Leaderboard V42026-04-15
Needs audit
16DeepSeek V4DeepSeek2026-04-2474.6%
Berkeley Function Calling Leaderboard V42026-04-15
Needs audit
17Llama 4 ScoutMeta2026-04-0574.1%
Berkeley Function Calling Leaderboard V42026-04-15
Needs audit
18DeepSeek R1DeepSeek2025-01-2072.4%
Berkeley Function Calling Leaderboard V42026-04-15
Needs audit
19Claude Haiku 4.5Anthropic2025-10-0171.8%
Berkeley Function Calling Leaderboard V42026-04-15
Needs audit
20Cohere Command ACohere2026-03-1570.6%
Berkeley Function Calling Leaderboard V42026-04-15
Needs audit
Method

What it covers

Primary source is the Berkeley Gorilla leaderboard. VerdictPal should store category and BFCL version with every score row.

Score ceiling

Where it breaks down

The leaderboard does not replace a product-level integration test with your schemas, auth model, error handling, and adversarial inputs.

Tools that report it

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.