Benchmarks / Agentic

τ³-Bench Banking

tau3-bench-banking

Dual-control agent-user simulation for banking tasks: 97 scenarios with knowledge retrieval, backend database state evaluation, pass@1. 14% weight in AA Intelligence Index v4.1.

What this does not measure
  • Real bank integrations — simulated environment with controlled APIs, not production core banking systems.
  • Regulatory compliance certification — passing τ³ does not mean a deployment meets SOC2 or banking regs.
  • Every customer-service domain — banking-specific policies and tools only.
Analysis

Why this benchmark is useful

Replaces τ²-Bench Telecom in AA Index v4.1 with harder, more realistic agentic scenarios that better separate frontier models.

Scope

Coverage map

Task family
Agentic
Format
97 banking scenarios; 5 repeats; dual-control simulation with tool use and database checks.
Scoring
Backend database state evaluation, pass@1.
Maintainer
Sierra / Artificial Analysis
Reading guide

How to read the scores

pass@1 on backend database state across 97 banking scenarios. Artificial Analysis now reports τ³ on a low scale (typical verified scores well below 60) — treat it as a live agentic separator, not a solved exam.

Blind spots

What it does not cover

  • Real bank integrations — simulated environment with controlled APIs, not production core banking systems.
  • Regulatory compliance certification — passing τ³ does not mean a deployment meets SOC2 or banking regs.
  • Every customer-service domain — banking-specific policies and tools only.
Scores

Evidence ledger

106 rows
106106 rows
51.34%best score
106source-checked
1sources
2026-09-01source date
106

api · api

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Qwen3.8 Max51.34% ·
  2. Grok 4.650.72% ·
  3. GLM-5.350.31% ·
  4. Qwen3.8 2.4T A95B49.07% ·
  5. Qwen3.8 27B48.04% ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Qwen3.8 MaxAlibaba2026-08-0351.34%
2026-09-01
Source-checked
2Grok 4.6xAI2026-08-1250.72%
2026-09-01
Source-checked
3GLM-5.3Zhipu2026-08-1450.31%
2026-09-01
Source-checked
4Qwen3.8 2.4T A95BAlibaba2026-08-1249.07%
2026-09-01
Source-checked
5Qwen3.8 27BAlibaba2026-08-1448.04%
2026-09-01
Source-checked
6GLM-5.3-FlashZhipu2026-08-2647.22%
2026-09-01
Source-checked
7Kimi K3Moonshot2026-07-1645.98%
2026-09-01
Source-checked
8Qwen3.8-Flash-NextAlibaba2026-08-2645.36%
2026-09-01
Source-checked
9Claude Opus 5 (Adaptive Reasoning, High Effort)Anthropic2026-07-2444.74%
2026-09-01
Source-checked
10GPT-5.6 SolOpenAI2026-07-0944.33%
2026-09-01
Source-checked
11Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)Anthropic2026-07-2443.3%
2026-09-01
Source-checked
12Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic2026-07-2442.06%
2026-09-01
Source-checked
13Grok 4.5xAI2026-07-0842.06%
2026-09-01
Source-checked
14Kimi K3 (low)Moonshot2026-07-1641.65%
2026-09-01
Source-checked
15DeepSeek V4 Flash Vision (Reasoning, Max Effort)DeepSeek2026-08-2141.03%
2026-09-01
Source-checked
16GPT-5.6 TerraOpenAI2026-07-0940.21%
2026-09-01
Source-checked
17DeepSeek V4 Pro 0813 (Reasoning, Max Effort)DeepSeek2026-08-1339.59%
2026-09-01
Source-checked
18GPT 5.4OpenAI2026-03-0539.59%
2026-09-01
Source-checked
19DeepSeek V4 Flash 0731 (Reasoning, Max Effort)DeepSeek2026-07-3139.38%
2026-09-01
Source-checked
20GPT 5.5OpenAI2026-04-2338.97%
2026-09-01
Source-checked
21Claude Opus 5 (Adaptive Reasoning, Medium Effort)Anthropic2026-07-2438.56%
2026-09-01
Source-checked
22Claude Fable 5Anthropic2026-06-0938.14%
2026-09-01
Source-checked
23Claude Sonnet 5Anthropic2026-06-3037.32%
2026-09-01
Source-checked
24GPT-5.5 (high)OpenAI2026-04-2336.7%
2026-09-01
Source-checked
25Agnes 2.5 Pro BetaOther2026-08-2635.67%
2026-09-01
Source-checked
26Motif 3Other2026-08-1235.26%
2026-09-01
Source-checked
27Muse Spark 1.2Meta2026-08-0534.85%
2026-09-01
Source-checked
28Claude Opus 4.7Anthropic2026-04-1634.64%
2026-09-01
Source-checked
29GLM-5.2Zhipu2026-06-1334.64%
2026-09-01
Source-checked
30Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-1734.43%
2026-09-01
Source-checked
31Claude Opus 4.8Anthropic2026-05-2834.23%
2026-09-01
Source-checked
32Gemini 3.7 FlashGoogle2026-08-1332.78%
2026-09-01
Source-checked
33Gemini 3.5 FlashGoogle2026-05-1932.16%
2026-09-01
Source-checked
34Muse Spark 1.1Meta2026-07-0931.75%
2026-09-01
Source-checked
35GPT-5.6 LunaOpenAI2026-07-0931.13%
2026-09-01
Source-checked
36Claude Opus 5 (Adaptive Reasoning, Low Effort)Anthropic2026-07-2430.31%
2026-09-01
Source-checked
37Gemini 3.6 Flash (high)Google2026-07-2129.9%
2026-09-01
Source-checked
38GPT-5.5 (medium)OpenAI2026-04-2329.9%
2026-09-01
Source-checked
39Inkling (xhigh)Other2026-07-1529.07%
2026-09-01
Source-checked
40GPT-5.4 NanoOpenAI2026-03-1727.42%
2026-09-01
Source-checked
41Ling 3.0 FlashOther2026-08-0427.22%
2026-09-01
Source-checked
42DeepSeek V4 Flash (Reasoning, High Effort)DeepSeek2026-04-2426.19%
2026-09-01
Source-checked
43DeepSeek V4 Pro (Reasoning, High Effort)DeepSeek2026-04-2426.19%
2026-09-01
Source-checked
44GPT-5.4 MiniOpenAI2026-03-1725.57%
2026-09-01
Source-checked
45Apodex 1.1Other2026-08-3025.15%
2026-09-01
Source-checked
46GPT-5.5 (low)OpenAI2026-04-2324.95%
2026-09-01
Source-checked
47Claude 4.5 Sonnet (Reasoning)Anthropic2025-09-2924.54%
2026-09-01
Source-checked
48Muse GlimmerMeta2026-08-1023.51%
2026-09-01
Source-checked
49Kimi K2.6Moonshot2026-04-2023.3%
2026-09-01
Source-checked
50Solar Pro 4Other2026-08-0623.3%
2026-09-01
Source-checked
51Hy3Other2026-07-0622.89%
2026-09-01
Source-checked
52G9v3-39A5BOther2026-08-0322.06%
2026-09-01
Source-checked
53GPT-5 (high)OpenAI2025-08-0722.06%
2026-09-01
Source-checked
54Solar Open2 250BOther2026-08-1221.65%
2026-09-01
Source-checked
55Gemini 3.1 Pro PreviewGoogle2026-02-1921.44%
2026-09-01
Source-checked
56Gemini 3 Flash Preview (Reasoning)Google2025-12-1720.82%
2026-09-01
Source-checked
57Ling 3.0 TinyOther2026-08-0620.82%
2026-09-01
Source-checked
58Qwen3.6 PlusAlibaba2026-04-0220.82%
2026-09-01
Source-checked
59Kimi K2.7 CodeMoonshot2026-06-1220.21%
2026-09-01
Source-checked
60Inkling SmallOther2026-07-3018.76%
2026-09-01
Source-checked
61Ring-2.6-1TOther2026-05-0817.94%
2026-09-01
Source-checked
62Gemini 3.5 Flash-LiteGoogle2026-07-2117.53%
2026-09-01
Source-checked
63Nex-N2-ProOther2026-06-0217.53%
2026-09-01
Source-checked
64Qwen 3.7 PlusAlibaba2026-06-0217.53%
2026-09-01
Source-checked
65GLM-5.2 (Non-reasoning)Zhipu2026-06-1616.7%
2026-09-01
Source-checked
66Qwen3.6 27B (Reasoning)Alibaba2026-04-2216.7%
2026-09-01
Source-checked
67A.X-K2Other2026-08-1216.08%
2026-09-01
Source-checked
68GPT-5.1OpenAI2025-11-1315.88%
2026-09-01
Source-checked
69MiniMax M3MiniMax2026-05-3115.26%
2026-09-01
Source-checked
70Qwen3.5 122B A10B (Reasoning)Alibaba2026-02-2415.26%
2026-09-01
Source-checked
71Mistral Medium 3.5Mistral2026-04-2915.05%
2026-09-01
Source-checked
72Gemma 4 31B ITGoogle2026-03-0114.85%
2026-09-01
Source-checked
73GPT-5.5 (Non-reasoning)OpenAI2026-04-2314.85%
2026-09-01
Source-checked
74Granite 4.2 30BOther2026-08-2514.43%
2026-09-01
Source-checked
75Kimi K2.5Moonshot2026-01-2714.23%
2026-09-01
Source-checked
76Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA2026-06-0414.23%
2026-09-01
Source-checked
77GLM-5.1Zhipu2026-04-0713.61%
2026-09-01
Source-checked
78Grok Build 0.1 0616xAI2026-06-1613.4%
2026-09-01
Source-checked
79Qwen3.5 397B A17B (Reasoning)Alibaba2026-02-1613.4%
2026-09-01
Source-checked
80LongCat-2.0Meituan2026-06-2913.2%
2026-09-01
Source-checked
81Agnes 2.5 Pro AlphaOther2026-07-2412.37%
2026-09-01
Source-checked
82Grok 4.3xAI2026-04-3012.37%
2026-09-01
Source-checked
83GLM-4.7 (Reasoning)Zhipu2025-12-2212.16%
2026-09-01
Source-checked
84Step 3.7 FlashStepFun2026-06-0111.96%
2026-09-01
Source-checked
85GPT-5.5 Instant (June 2026)OpenAI2026-06-2511.75%
2026-09-01
Source-checked
86Qwen3.7 MaxAlibaba2026-05-1911.75%
2026-09-01
Source-checked
87K-EXAONE 2.0 0803Other2026-08-1211.55%
2026-09-01
Source-checked
88MiMo-V2.5-ProXiaomi2026-04-229.9%
2026-09-01
Source-checked
89MiniMax M2.7MiniMax2026-04-159.9%
2026-09-01
Source-checked
90Gemini 3.1 Flash LiteGoogle2026-05-079.69%
2026-09-01
Source-checked
91Qwen3.6 27B (Non-reasoning)Alibaba2026-04-229.28%
2026-09-01
Source-checked
92Qwen3.6 35B A3B (Reasoning)Alibaba2026-04-169.28%
2026-09-01
Source-checked
93Nemotron 3.5 LightningNVIDIA2026-08-118.87%
2026-09-01
Source-checked
94MiMo-V2.5Xiaomi2026-04-228.66%
2026-09-01
Source-checked
95Grok 4.3 (Non-reasoning)xAI2026-04-308.04%
2026-09-01
Source-checked
96Granite 4.2 8BOther2026-08-257.63%
2026-09-01
Source-checked
97North Mini CodeCohere2026-06-096.39%
2026-09-01
Source-checked
98Command ACohere2026-03-155.98%
2026-09-01
Source-checked
99Mistral Large 3Mistral2025-12-025.77%
2026-09-01
Source-checked
100Granite 4.2 3BOther2026-08-255.57%
2026-09-01
Source-checked
101HyperNova 60B 2605Multiverse2026-05-065.57%
2026-09-01
Source-checked
102Qwen3.6 35B A3B (Non-reasoning)Alibaba2026-04-165.36%
2026-09-01
Source-checked
103KAT Coder Pro V2Other2026-03-274.54%
2026-09-01
Source-checked
104Llama 4 MaverickMeta2026-04-053.71%
2026-09-01
Source-checked
105Llama 4 ScoutMeta2026-04-053.3%
2026-09-01
Source-checked
106Ling 2.6 FlashOther2026-04-212.89%
2026-09-01
Source-checked
Method

What it covers

Live — AA τ³-Bench Banking (Index v4.1, 14% weight) uses a harder dual-control scale than retired τ². VerdictPal maps only current `tau_banking` / `tau3_banking` rows onto this page; older τ² high-scale receipts are not mixed in.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Tools that report it

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does τ³-Bench Banking measure?

Dual-control agent-user simulation for banking tasks: 97 scenarios with knowledge retrieval, backend database state evaluation, pass@1. 14% weight in AA Intelligence Index v4.1.

What does a high τ³-Bench Banking score not prove?

A strong τ³-Bench Banking result says nothing about:

  • Real bank integrations — simulated environment with controlled APIs, not production core banking systems.
  • Regulatory compliance certification — passing τ³ does not mean a deployment meets SOC2 or banking regs.
  • Every customer-service domain — banking-specific policies and tools only.

How is τ³-Bench Banking scored?

Backend database state evaluation, pass@1. Task format: 97 banking scenarios; 5 repeats; dual-control simulation with tool use and database checks.

Is τ³-Bench Banking saturated?

τ³-Bench Banking is currently marked Active in the atlas.

Can τ³-Bench Banking results be contaminated by training data?

Contamination risk for τ³-Bench Banking is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.