Benchmarks / Performance

AA output speed

Output tokens per second

Median output tokens per second after the first chunk arrives, measured on each model's default API provider with a ~1k-token input prompt (AA medium workload).

Artificial AnalysisPerformanceFlagshipActiveUnknown contamination riskSince 2026
What this does not measure
  • IDE or chat-app latency — only first-party API routes AA tracks as the default provider.
  • Quality or reasoning depth — speed-only slice. A fast model can still fail hard tasks.
Analysis

Why this benchmark is useful

Editorial brief pendingWe publish the methodology and ledger first; benchmark-specific analysis ships after desk review.

Scope

Coverage map

Task family
Performance
Format
Streaming generation after first token; median over the past 72 hours on AA's harness.
Scoring
Higher tokens/s is faster. AA uses OpenAI tiktoken o200k_base for token counts.
Maintainer
Artificial Analysis
Reading guide

How to read the scores

Reading guide pending. Use task format, scoring method, and source dates in the ledger until the desk brief ships.

Blind spots

What it does not cover

  • IDE or chat-app latency — only first-party API routes AA tracks as the default provider.
  • Quality or reasoning depth — speed-only slice. A fast model can still fail hard tasks.
Scores

Evidence ledger

69 rows
6969 rows
401 t/sbest score
69source-checked
1sources
2026-07-21source date
69

api · api

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Step 3.7 Flash401 t/s ·
  2. HyperNova 60B 2605381 t/s ·
  3. LFM2.5-8B-A1B343 t/s ·
  4. Gemini 3.1 Flash Lite315 t/s ·
  5. Gemini 3.5 Flash (medium)299 t/s ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Step 3.7 FlashStepFun2026-06-01401 t/s
2026-07-21
Source-checked
2HyperNova 60B 2605Multiverse2026-05-06381 t/s
2026-07-21
Source-checked
3LFM2.5-8B-A1BLiquid2026-06-08343 t/s
2026-07-21
Source-checked
4Gemini 3.1 Flash LiteGoogle2026-05-07315 t/s
2026-07-21
Source-checked
5Gemini 3.5 Flash (medium)Google2026-05-19299 t/s
2026-07-21
Source-checked
6Gemini 3.5 FlashGoogle2026-05-19287 t/s
2026-07-21
Source-checked
7Grok 4.20 ReasoningxAI2026-03-05227 t/s
2026-07-21
Source-checked
8Gemini 3 Flash PreviewGoogle2025-12-17218 t/s
2026-07-21
Source-checked
9Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA2026-06-04215 t/s
2026-07-21
Source-checked
10Qwen3.7 MaxAlibaba2026-05-19207 t/s
2026-07-21
Source-checked
11GPT-5.6 LunaOpenAI2026-07-09203 t/s
2026-07-21
Source-checked
12GLM-5.2Zhipu2026-06-13196 t/s
2026-07-21
Source-checked
13GPT-5.4 MiniOpenAI2026-03-17180 t/s
2026-07-21
Source-checked
14GPT-5.4 NanoOpenAI2026-03-17163 t/s
2026-07-21
Source-checked
15GPT-5.6 TerraOpenAI2026-07-09157 t/s
2026-07-21
Source-checked
16OpenAI o3OpenAI2025-04-16155 t/s
2026-07-21
Source-checked
17GPT 5.4OpenAI2026-03-05152 t/s
2026-07-21
Source-checked
18Gemini 3.1 Pro PreviewGoogle2026-02-19136 t/s
2026-07-21
Source-checked
19Gemma 4 12B (Reasoning)Google2026-06-07129 t/s
2026-07-21
Source-checked
20GPT-5.3 Codex (xhigh)OpenAI2026-02-05127 t/s
2026-07-21
Source-checked
21Grok 4.3xAI2026-04-30124 t/s
2026-07-21
Source-checked
22Muse Spark 1.1Meta2026-07-09122 t/s
2026-07-21
Source-checked
23Gemma 4 12B (Non-reasoning)Google2026-06-10119 t/s
2026-07-21
Source-checked
24Llama 4 MaverickMeta2026-04-05110 t/s
2026-07-21
Source-checked
25GPT-5.1OpenAI2025-11-13110 t/s
2026-07-21
Source-checked
26MiniMax M3MiniMax2026-05-3197 t/s
2026-07-21
Source-checked
27North Mini CodeCohere2026-06-0996 t/s
2026-07-21
Source-checked
28GPT 5.5OpenAI2026-04-2390 t/s
2026-07-21
Source-checked
29Claude Sonnet 5Anthropic2026-06-3088 t/s
2026-07-21
Source-checked
30GPT-5.2OpenAI2025-12-1187 t/s
2026-07-21
Source-checked
31MiniMax M2.5MiniMax2026-04-0187 t/s
2026-07-21
Source-checked
32GLM-5.1Zhipu2026-04-0785 t/s
2026-07-21
Source-checked
33Grok 4.5xAI2026-07-0877 t/s
2026-07-21
Source-checked
34GPT-5.5 (high)OpenAI2026-04-2377 t/s
2026-07-21
Source-checked
35Llama 4 ScoutMeta2026-04-0576 t/s
2026-07-21
Source-checked
36GPT-5.5 (medium)OpenAI2026-04-2376 t/s
2026-07-21
Source-checked
37GPT-5.5 (low)OpenAI2026-04-2376 t/s
2026-07-21
Source-checked
38DeepSeek V4 Pro (Reasoning, Max Effort)DeepSeek2026-04-2473 t/s
2026-07-21
Source-checked
39Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-1772 t/s
2026-07-21
Source-checked
40Claude Fable 5Anthropic2026-06-0970 t/s
2026-07-21
Source-checked
41Claude Opus 4.5Anthropic2025-11-2468 t/s
2026-07-21
Source-checked
42GPT-5.6 SolOpenAI2026-07-0967 t/s
2026-07-21
Source-checked
43Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-0564 t/s
2026-07-21
Source-checked
44Mistral Large 3Mistral2025-12-0264 t/s
2026-07-21
Source-checked
45Claude Opus 4.8Anthropic2026-05-2863 t/s
2026-07-21
Source-checked
46Claude Opus 4.7Anthropic2026-04-1661 t/s
2026-07-21
Source-checked
47Command ACohere2026-03-1561 t/s
2026-07-21
Source-checked
48Claude Sonnet 4.6Anthropic2026-02-1760 t/s
2026-07-21
Source-checked
49MiniMax M2.7MiniMax2026-04-1559 t/s
2026-07-21
Source-checked
50MiMo-V2.5-ProXiaomi2026-04-2258 t/s
2026-07-21
Source-checked
51Claude Opus 4.6Anthropic2026-02-0557 t/s
2026-07-21
Source-checked
52Kimi K2.6Moonshot2026-04-2056 t/s
2026-07-21
Source-checked
53Qwen 3.7 PlusAlibaba2026-06-0253 t/s
2026-07-21
Source-checked
54Claude Opus 4.7 (Non-reasoning, High Effort)Anthropic2026-04-1650 t/s
2026-07-21
Source-checked
55Kimi K2.5Moonshot2026-01-2750 t/s
2026-07-21
Source-checked
56Kimi K2.7 CodeMoonshot2026-06-1247 t/s
2026-07-21
Source-checked
57Kimi K3Moonshot2026-07-1638 t/s
2026-07-21
Source-checked
58Gemma 4 31B ITGoogle2026-03-0136 t/s
2026-07-21
Source-checked
59Claude 3 OpusAnthropic2024-03-040 t/s
2026-07-21
Source-checked
60DeepSeek Coder V2DeepSeek2024-05-150 t/s
2026-07-21
Source-checked
61DeepSeek R1DeepSeek2025-01-200 t/s
2026-07-21
Source-checked
62Gemini 1.0 UltraGoogle2024-02-080 t/s
2026-07-21
Source-checked
63Gemini 3 ProGoogle2025-11-180 t/s
2026-07-21
Source-checked
64GPT-5.4 ProOpenAI2026-03-050 t/s
2026-07-21
Source-checked
65GPT-5.5 ProOpenAI2026-04-230 t/s
2026-07-21
Source-checked
66LongCat-2.0Meituan2026-06-290 t/s
2026-07-21
Source-checked
67MiniCPM5-1B (Reasoning)OpenBMB2026-06-040 t/s
2026-07-21
Source-checked
68Muse SparkOther2026-01-010 t/s
2026-07-21
Source-checked
69OpenAI o1OpenAI2024-09-120 t/s
2026-07-21
Source-checked
Method

What it covers

Data-quality note pending. Every ledger row still carries source URL, source date, and ingest timestamp.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Tools that report it

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.