Benchmarks / Performance

AA output speed

Output tokens per second

Median output tokens per second after the first chunk arrives, measured on each model's default API provider with a ~1k-token input prompt (AA medium workload).

What this does not measure
  • IDE or chat-app latency — only first-party API routes AA tracks as the default provider.
  • Quality or reasoning depth — speed-only slice. A fast model can still fail hard tasks.
Analysis

Why this benchmark is useful

When you care how fast tokens stream after the first chunk — IDE feel, chat responsiveness, batch throughput — not whether the answer is right.

Scope

Coverage map

Task family
Performance
Format
Streaming generation after first token; median over the past 72 hours on AA's harness.
Scoring
Higher tokens/s is faster. AA uses OpenAI tiktoken o200k_base for token counts.
Maintainer
Artificial Analysis
Reading guide

How to read the scores

Higher median tokens/s is faster. AA pins each model's default API provider and a ~1k-token medium workload; IDE or alternate providers can diverge.

Blind spots

What it does not cover

  • IDE or chat-app latency — only first-party API routes AA tracks as the default provider.
  • Quality or reasoning depth — speed-only slice. A fast model can still fail hard tasks.
Scores

Evidence ledger

187 rows
187187 rows
373 t/sbest score
187source-checked
1sources
2026-09-05to 2026-09-01
187

api · api

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Muse Spark 1.3190 t/s ·
  2. GPT-6 Astra64 t/s ·
  3. Ling 3.0 Flash373 t/s ·
  4. HyperNova 60B 2605351 t/s ·
  5. LFM2.5-8B-A1B346 t/s ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Ling 3.0 FlashOther2026-08-04373 t/s
2026-09-01
Source-checked
2HyperNova 60B 2605Multiverse2026-05-06351 t/s
2026-09-01
Source-checked
3LFM2.5-8B-A1BLiquid2026-06-08346 t/s
2026-09-01
Source-checked
4Nemotron 3 Nano Omni 30B A3B ReasoningNVIDIA2026-04-29332 t/s
2026-09-01
Source-checked
5Gemini 3.5 Flash-LiteGoogle2026-07-21316 t/s
2026-09-01
Source-checked
6Gemini 3.7 FlashGoogle2026-08-13307 t/s
2026-09-01
Source-checked
7Nemotron 3.5 LightningNVIDIA2026-08-11290 t/s
2026-09-01
Source-checked
8Granite 4.2 3BOther2026-08-25234 t/s
2026-09-01
Source-checked
9Gemini 3.5 Flash (medium)Google2026-05-19232 t/s
2026-09-01
Source-checked
10Command ACohere2026-03-15228 t/s
2026-09-01
Source-checked
11Gemini 3.5 Flash (minimal)Google2026-05-19217 t/s
2026-09-01
Source-checked
12Ling 3.0 TinyOther2026-08-06192 t/s
2026-09-01
Source-checked
13Muse Spark 1.3Meta2026-09-02190 t/s
2026-09-05
Source-checked
14Agnes 2.5 Pro AlphaOther2026-07-24174 t/s
2026-09-01
Source-checked
15Gemini 3.6 Flash (high)Google2026-07-21161 t/s
2026-09-01
Source-checked
16Agnes 2.5 Pro BetaOther2026-08-26152 t/s
2026-09-01
Source-checked
17Granite 4.2 8BOther2026-08-25152 t/s
2026-09-01
Source-checked
18Qwen3.6 35B A3B (Non-reasoning)Alibaba2026-04-16139 t/s
2026-09-01
Source-checked
19Nex-N2-ProOther2026-06-02135 t/s
2026-09-01
Source-checked
20Mistral Medium 3.5Mistral2026-04-29135 t/s
2026-09-01
Source-checked
21Qwen3.5 122B A10B (Reasoning)Alibaba2026-02-24134 t/s
2026-09-01
Source-checked
22OpenAI o3OpenAI2025-04-16134 t/s
2026-09-01
Source-checked
23Llama 4 ScoutMeta2026-04-05133 t/s
2026-09-01
Source-checked
24Gemma 4 12B (Reasoning)Google2026-06-07131 t/s
2026-09-01
Source-checked
25Gemma 4 12B (Non-reasoning)Google2026-06-10130 t/s
2026-09-01
Source-checked
26Qwen3.6 35B A3B (Reasoning)Alibaba2026-04-16130 t/s
2026-09-01
Source-checked
27GPT-5.5 Instant (June 2026)OpenAI2026-06-25127 t/s
2026-09-01
Source-checked
28Ring-2.6-1TOther2026-05-08125 t/s
2026-09-01
Source-checked
29GPT-5.6 LunaOpenAI2026-07-09123 t/s
2026-09-01
Source-checked
30Grok 4.3 (medium)xAI2026-04-30121 t/s
2026-09-01
Source-checked
31GPT-5.3 Codex (xhigh)OpenAI2026-02-05117 t/s
2026-09-01
Source-checked
32MiniMax M3MiniMax2026-05-31114 t/s
2026-09-01
Source-checked
33Gemini 3.1 Pro PreviewGoogle2026-02-19113 t/s
2026-09-01
Source-checked
34DeepSeek V4 Flash Vision (Reasoning, Max Effort)DeepSeek2026-08-21111 t/s
2026-09-01
Source-checked
35KAT Coder Pro V2Other2026-03-27108 t/s
2026-09-01
Source-checked
36GPT-5.6 TerraOpenAI2026-07-09108 t/s
2026-09-01
Source-checked
37DeepSeek V4 Flash 0731 (Reasoning, Max Effort)DeepSeek2026-07-31106 t/s
2026-09-01
Source-checked
38Grok 4.3 (low)xAI2026-04-30104 t/s
2026-09-01
Source-checked
39Muse GlimmerMeta2026-08-10101 t/s
2026-09-01
Source-checked
40GLM-5.2 (Non-reasoning)Zhipu2026-06-16100 t/s
2026-09-01
Source-checked
41Hy3Other2026-07-0698 t/s
2026-09-01
Source-checked
42Grok 4.3 (Non-reasoning)xAI2026-04-3098 t/s
2026-09-01
Source-checked
43Step 3.7 FlashStepFun2026-06-0195 t/s
2026-09-01
Source-checked
44North Mini CodeCohere2026-06-0991 t/s
2026-09-01
Source-checked
45Llama 4 MaverickMeta2026-04-0590 t/s
2026-09-01
Source-checked
46Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA2026-06-0487 t/s
2026-09-01
Source-checked
47Qwen3.8-Flash-NextAlibaba2026-08-2687 t/s
2026-09-01
Source-checked
48Qwen3.5 397B A17B (Reasoning)Alibaba2026-02-1685 t/s
2026-09-01
Source-checked
49Qwen3.5 397B A17B (Non-reasoning)Alibaba2026-02-1681 t/s
2026-09-01
Source-checked
50Claude Sonnet 5Anthropic2026-06-3080 t/s
2026-09-01
Source-checked
51Mistral Large 3Mistral2025-12-0278 t/s
2026-09-01
Source-checked
52Granite 4.2 30BOther2026-08-2577 t/s
2026-09-01
Source-checked
53GPT-5.6 SolOpenAI2026-07-0977 t/s
2026-09-01
Source-checked
54Inkling (xhigh)Other2026-07-1574 t/s
2026-09-01
Source-checked
55GLM-5.3Zhipu2026-08-1473 t/s
2026-09-01
Source-checked
56GLM-5.2Zhipu2026-06-1369 t/s
2026-09-01
Source-checked
57GPT-6 AstraOpenAI2026-09-0364 t/s
2026-09-05
Source-checked
58MiMo-V2.5Xiaomi2026-04-2258 t/s
2026-09-01
Source-checked
59Claude Fable 5Anthropic2026-06-0958 t/s
2026-09-01
Source-checked
60Qwen 3.7 PlusAlibaba2026-06-0256 t/s
2026-09-01
Source-checked
61Qwen3.6 27B (Non-reasoning)Alibaba2026-04-2255 t/s
2026-09-01
Source-checked
62Qwen3.6 27B (Reasoning)Alibaba2026-04-2253 t/s
2026-09-01
Source-checked
63Solar Pro 4Other2026-08-0653 t/s
2026-09-01
Source-checked
64Grok 4.6xAI2026-08-1251 t/s
2026-09-01
Source-checked
65DeepSeek V4 Pro 0813 (Reasoning, Max Effort)DeepSeek2026-08-1351 t/s
2026-09-01
Source-checked
66DeepSeek V4 Pro (Reasoning, High Effort)DeepSeek2026-04-2450 t/s
2026-09-01
Source-checked
67DeepSeek V4 Pro (Non-reasoning)DeepSeek2026-04-2450 t/s
2026-09-01
Source-checked
68Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic2026-07-2449 t/s
2026-09-01
Source-checked
69Grok 4.5xAI2026-07-0848 t/s
2026-09-01
Source-checked
70Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)Anthropic2026-07-2448 t/s
2026-09-01
Source-checked
71Claude Opus 5 (Adaptive Reasoning, High Effort)Anthropic2026-07-2447 t/s
2026-09-01
Source-checked
72Claude Opus 5 (Adaptive Reasoning, Low Effort)Anthropic2026-07-2447 t/s
2026-09-01
Source-checked
73Inkling SmallOther2026-07-3047 t/s
2026-09-01
Source-checked
74Claude Opus 5 (Adaptive Reasoning, Medium Effort)Anthropic2026-07-2446 t/s
2026-09-01
Source-checked
75Qwen3.8 27BAlibaba2026-08-1445 t/s
2026-09-01
Source-checked
76Claude Sonnet 4.6 (Non-reasoning, Low Effort)Anthropic2026-02-1745 t/s
2026-09-01
Source-checked
77Kimi K2.7 CodeMoonshot2026-06-1244 t/s
2026-09-01
Source-checked
78GLM-5.3-FlashZhipu2026-08-2642 t/s
2026-09-01
Source-checked
79Qwen3.8 MaxAlibaba2026-08-0340 t/s
2026-09-01
Source-checked
80Qwen3.8 2.4T A95BAlibaba2026-08-1240 t/s
2026-09-01
Source-checked
81Kimi K3Moonshot2026-07-1640 t/s
2026-09-01
Source-checked
82LongCat-2.0Meituan2026-06-2937 t/s
2026-09-01
Source-checked
83MiMo-V2.5-ProXiaomi2026-04-2237 t/s
2026-09-01
Source-checked
84Kimi K3 (low)Moonshot2026-07-1636 t/s
2026-09-01
Source-checked
85Gemma 4 31B ITGoogle2026-03-0136 t/s
2026-09-01
Source-checked
86MiMo-V2.5-Pro (Non-reasoning)Xiaomi2026-04-2230 t/s
2026-09-01
Source-checked
87A.X-K2Other2026-08-120 t/s
2026-09-01
Source-checked
88Apodex 1.1Other2026-08-300 t/s
2026-09-01
Source-checked
89Claude 3 OpusAnthropic2024-03-040 t/s
2026-09-01
Source-checked
90Claude 4 Opus (Reasoning)Anthropic2025-05-220 t/s
2026-09-01
Source-checked
91Claude 4.1 Opus (Reasoning)Anthropic2025-08-050 t/s
2026-09-01
Source-checked
92Claude 4.5 Sonnet (Reasoning)Anthropic2025-09-290 t/s
2026-09-01
Source-checked
93Claude Opus 4.5Anthropic2025-11-240 t/s
2026-09-01
Source-checked
94Claude Opus 4.5 (Reasoning)Anthropic2025-11-240 t/s
2026-09-01
Source-checked
95Claude Opus 4.6Anthropic2026-02-050 t/s
2026-09-01
Source-checked
96Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-050 t/s
2026-09-01
Source-checked
97Claude Opus 4.7Anthropic2026-04-160 t/s
2026-09-01
Source-checked
98Claude Opus 4.7 (Non-reasoning, High Effort)Anthropic2026-04-160 t/s
2026-09-01
Source-checked
99Claude Opus 4.8Anthropic2026-05-280 t/s
2026-09-01
Source-checked
100Claude Sonnet 4.6Anthropic2026-02-170 t/s
2026-09-01
Source-checked
101Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-170 t/s
2026-09-01
Source-checked
102DeepSeek Coder V2DeepSeek2024-05-150 t/s
2026-09-01
Source-checked
103DeepSeek R1DeepSeek2025-01-200 t/s
2026-09-01
Source-checked
104DeepSeek V3.2 (Reasoning)DeepSeek2025-12-010 t/s
2026-09-01
Source-checked
105DeepSeek V4 Flash (Non-reasoning)DeepSeek2026-04-240 t/s
2026-09-01
Source-checked
106DeepSeek V4 Flash (Reasoning, High Effort)DeepSeek2026-04-240 t/s
2026-09-01
Source-checked
107DiffusionGemma 26B A4BGoogle2026-06-100 t/s
2026-09-01
Source-checked
108EXAONE 4.5 33BOther2026-04-090 t/s
2026-09-01
Source-checked
109G9v3-39A5BOther2026-08-030 t/s
2026-09-01
Source-checked
110G9v3-3BOther2026-07-230 t/s
2026-09-01
Source-checked
111Gemini 1.0 UltraGoogle2024-02-080 t/s
2026-09-01
Source-checked
112Gemini 3 Flash PreviewGoogle2025-12-170 t/s
2026-09-01
Source-checked
113Gemini 3 Flash Preview (Reasoning)Google2025-12-170 t/s
2026-09-01
Source-checked
114Gemini 3 ProGoogle2025-11-180 t/s
2026-09-01
Source-checked
115Gemini 3 Pro Preview (low)Google2025-11-180 t/s
2026-09-01
Source-checked
116Gemini 3.1 Flash LiteGoogle2026-05-070 t/s
2026-09-01
Source-checked
117Gemini 3.5 FlashGoogle2026-05-190 t/s
2026-09-01
Source-checked
118GLM 5V Turbo (Reasoning)Zhipu2026-04-010 t/s
2026-09-01
Source-checked
119GLM-4.7 (Reasoning)Zhipu2025-12-220 t/s
2026-09-01
Source-checked
120GLM-5 (Non-reasoning)Zhipu2026-02-110 t/s
2026-09-01
Source-checked
121GLM-5 (Reasoning)Zhipu2026-02-110 t/s
2026-09-01
Source-checked
122GLM-5-TurboZhipu2026-03-150 t/s
2026-09-01
Source-checked
123GLM-5.1Zhipu2026-04-070 t/s
2026-09-01
Source-checked
124GPT 5.4OpenAI2026-03-050 t/s
2026-09-01
Source-checked
125GPT 5.5OpenAI2026-04-230 t/s
2026-09-01
Source-checked
126GPT-5 (high)OpenAI2025-08-070 t/s
2026-09-01
Source-checked
127GPT-5 (low)OpenAI2025-08-070 t/s
2026-09-01
Source-checked
128GPT-5 (medium)OpenAI2025-08-070 t/s
2026-09-01
Source-checked
129GPT-5 Codex (high)OpenAI2025-09-230 t/s
2026-09-01
Source-checked
130GPT-5.1OpenAI2025-11-130 t/s
2026-09-01
Source-checked
131GPT-5.1 Codex (high)OpenAI2025-11-130 t/s
2026-09-01
Source-checked
132GPT-5.2OpenAI2025-12-110 t/s
2026-09-01
Source-checked
133GPT-5.2 (medium)OpenAI2025-12-110 t/s
2026-09-01
Source-checked
134GPT-5.2 Codex (xhigh)OpenAI2025-12-110 t/s
2026-09-01
Source-checked
135GPT-5.4 (low)OpenAI2026-03-050 t/s
2026-09-01
Source-checked
136GPT-5.4 MiniOpenAI2026-03-170 t/s
2026-09-01
Source-checked
137GPT-5.4 NanoOpenAI2026-03-170 t/s
2026-09-01
Source-checked
138GPT-5.4 ProOpenAI2026-03-050 t/s
2026-09-01
Source-checked
139GPT-5.5 (high)OpenAI2026-04-230 t/s
2026-09-01
Source-checked
140GPT-5.5 (low)OpenAI2026-04-230 t/s
2026-09-01
Source-checked
141GPT-5.5 (medium)OpenAI2026-04-230 t/s
2026-09-01
Source-checked
142GPT-5.5 (Non-reasoning)OpenAI2026-04-230 t/s
2026-09-01
Source-checked
143GPT-5.5 Instant (May 2026)OpenAI2026-05-050 t/s
2026-09-01
Source-checked
144GPT-5.5 ProOpenAI2026-04-230 t/s
2026-09-01
Source-checked
145Granite 4.1 30BOther2026-04-290 t/s
2026-09-01
Source-checked
146Grok 4xAI2025-07-100 t/s
2026-09-01
Source-checked
147Grok 4.20 0309 (Reasoning)xAI2026-03-100 t/s
2026-09-01
Source-checked
148Grok 4.20 0309 v2 (Non-reasoning)xAI2026-04-070 t/s
2026-09-01
Source-checked
149Grok 4.20 ReasoningxAI2026-03-050 t/s
2026-09-01
Source-checked
150Grok 4.3xAI2026-04-300 t/s
2026-09-01
Source-checked
151Grok Build 0.1 0616xAI2026-06-160 t/s
2026-09-01
Source-checked
152Hy3-preview (Non-reasoning)Other2026-04-230 t/s
2026-09-01
Source-checked
153Hy3-preview (Reasoning)Other2026-04-230 t/s
2026-09-01
Source-checked
154JT-35B-FlashOther2026-05-140 t/s
2026-09-01
Source-checked
155JT-4.1 Flash 236B A21BOther2026-07-090 t/s
2026-09-01
Source-checked
156JT-MINIOther2026-04-150 t/s
2026-09-01
Source-checked
157K-EXAONE 2.0 0803Other2026-08-120 t/s
2026-09-01
Source-checked
158Kimi K2 ThinkingMoonshot2025-11-060 t/s
2026-09-01
Source-checked
159Kimi K2.5Moonshot2026-01-270 t/s
2026-09-01
Source-checked
160Kimi K2.6Moonshot2026-04-200 t/s
2026-09-01
Source-checked
161Kimi K2.6 (Non-reasoning)Other2026-04-200 t/s
2026-09-01
Source-checked
162Ling 2.6 FlashOther2026-04-210 t/s
2026-09-01
Source-checked
163Ling-2.6-1TOther2026-04-230 t/s
2026-09-01
Source-checked
164MiMo-V2-Flash (Feb 2026)Xiaomi2025-12-160 t/s
2026-09-01
Source-checked
165MiMo-V2-Flash (Reasoning)Xiaomi2025-12-160 t/s
2026-09-01
Source-checked
166MiMo-V2-OmniXiaomi2026-03-190 t/s
2026-09-01
Source-checked
167MiMo-V2-Omni-0327Xiaomi2026-03-270 t/s
2026-09-01
Source-checked
168MiMo-V2-ProXiaomi2026-03-180 t/s
2026-09-01
Source-checked
169MiniCPM-V 4.6 1.3BOpenBMB2026-05-110 t/s
2026-09-01
Source-checked
170MiniCPM5-1B (Non-reasoning)OpenBMB2026-05-250 t/s
2026-09-01
Source-checked
171MiniCPM5-1B (Reasoning)OpenBMB2026-06-040 t/s
2026-09-01
Source-checked
172MiniMax M2.5MiniMax2026-04-010 t/s
2026-09-01
Source-checked
173MiniMax M2.7MiniMax2026-04-150 t/s
2026-09-01
Source-checked
174MiniMax-M2.1MiniMax2025-12-230 t/s
2026-09-01
Source-checked
175Motif 3Other2026-08-120 t/s
2026-09-01
Source-checked
176Motif 3 (Beta)Other2026-07-210 t/s
2026-09-01
Source-checked
177Muse SparkOther2026-01-010 t/s
2026-09-01
Source-checked
178Muse Spark 1.1Meta2026-07-090 t/s
2026-09-01
Source-checked
179Muse Spark 1.2Meta2026-08-050 t/s
2026-09-01
Source-checked
180o3-proOpenAI2025-06-100 t/s
2026-09-01
Source-checked
181OpenAI o1OpenAI2024-09-120 t/s
2026-09-01
Source-checked
182Qwen3 Max ThinkingAlibaba2026-01-260 t/s
2026-09-01
Source-checked
183Qwen3.5 27B (Reasoning)Alibaba2026-02-240 t/s
2026-09-01
Source-checked
184Qwen3.6 Max PreviewAlibaba2026-04-200 t/s
2026-09-01
Source-checked
185Qwen3.6 PlusAlibaba2026-04-020 t/s
2026-09-01
Source-checked
186Qwen3.7 MaxAlibaba2026-05-190 t/s
2026-09-01
Source-checked
187Solar Open2 250BOther2026-08-120 t/s
2026-09-01
Source-checked
Method

What it covers

This benchmark sits in the performance family. It uses Streaming generation after first token; median over the past 72 hours on AA's harness. The score should travel with its task format, scoring method, source date, and benchmark version.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Tools that report it

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does AA output speed measure?

Median output tokens per second after the first chunk arrives, measured on each model's default API provider with a ~1k-token input prompt (AA medium workload).

What does a high AA output speed score not prove?

A strong AA output speed result says nothing about:

  • IDE or chat-app latency — only first-party API routes AA tracks as the default provider.
  • Quality or reasoning depth — speed-only slice. A fast model can still fail hard tasks.

How is AA output speed scored?

Higher tokens/s is faster. AA uses OpenAI tiktoken o200k_base for token counts. Task format: Streaming generation after first token; median over the past 72 hours on AA's harness.

Is AA output speed saturated?

AA output speed is currently marked Active in the atlas.

Can AA output speed results be contaminated by training data?

Contamination risk for AA output speed is graded Unknown contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.