Models / Alibaba

Qwen3.8 Max

Qwen · Released 2026-08-03

Alibaba's August 2026 Max-tier flagship. Artificial Analysis Intelligence Index 58.1. API list price $2 / $6 per 1M tokens on the AA snapshot.

Current Alibaba Cloud Max model for long-context and agentic API work. Prefer sourced ledger rows over vendor eval tables; AA scores sync from the public snapshot.

#6 of 166 on AA Index · current snapshot
58 AA Index · 10 benchmark rows · 1 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite66
Context1M
Input / 1M$2.0
Output / 1M$6.0
Knowledge cutoffunpublished
Output speed40 t/s
TTFT (default API)1.46s
Statusflagship

Family profile

Best published score in each covered benchmark family.

77/100 avg
Reasoning
93
Coding
53
Agentic
81
Multimodal
82
Long context
74

5 tested benchmark families

Editor's note
Recorded from the 2026-08-12 AA snapshot (slug qwen3-8-max, index 58.1) plus Alibaba Cloud launch copy. Editorial depth and failure modes still need a desk pass — scores below are ledger-backed, not hands-on claims. Vendor PaperBench/OSWorld figures stay off the ledger until they appear in a mapped source.
Benchmark placements

Where Qwen3.8 Max places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
τ³-Bench BankingAgentic#1/ 10651Artificial Analysis2026-09-01
MMMU-ProMultimodal#6/ 882Artificial Analysis — MMMU-Pro2026-09-01
Artificial Analysis Intelligence IndexReasoning#11/ 18458Artificial Analysis2026-09-01
GPQA DiamondReasoning#16/ 18593Artificial Analysis2026-09-01
Humanity's Last ExamReasoning#17/ 18043Artificial Analysis2026-09-01
Terminal-BenchAgentic#20/ 18481Artificial Analysis2026-09-01
SciCodeCoding#28/ 18053Artificial Analysis2026-09-01
AA time to first tokenPerformance#45/ 185Artificial Analysis2026-09-01
AA-LCRLong context#56/ 17974Artificial Analysis2026-09-01
AA output speedPerformance#77/ 185Artificial Analysis2026-09-01
Family context

The three highest-scoring models with pages in each capability family. Where Qwen3.8 Max shows up, it's highlighted.

Newest receipts

Benchmark rows added to the public ledger for Qwen3.8 Max in the last 120 days. Older rows live in the full table below.

10 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Reasoning93
Coding53
Agentic81
Multimodal82
Long context74
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMetaMiniMaxxAIZhipuXiaomiOtherMeituanCohereNVIDIA

Full benchmark ledger

Every catalog benchmark for Qwen3.8 Max. Scores link to the original source; gaps mean no public row exists yet.

10 sourced rows

10 of 37 catalog benchmarks have a sourced row for Qwen3.8 Max.

Reasoning

Coding

Agentic

Multimodal

Long context

Performance

Changelog
  1. Qwen3.8 Max appears on Artificial Analysis (index 58.1, $2 / $6 per 1M tokens).
Sources