Models / Alibaba

Qwen3 Max Thinking

Qwen3 Max · Released 2026-01-26

AA Intelligence Index 32.5. Listed on Artificial Analysis with API pricing $0/$0 per 1M tokens. Benchmark rows sync from the AA snapshot — see the ledger on this page.

Frontier model tracked on Artificial Analysis. Use the benchmark ledger below for sourced scores; we do not infer numbers beyond published rows.

#112 of 166 on AA Index · current snapshot
33 AA Index · 9 benchmark rows · 1 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite51
ContextUnlisted
Input / 1MFree
Output / 1MFree
Knowledge cutoffunpublished
Statuslightweight

Family profile

Best published score in each covered benchmark family.

56/100 avg
Reasoning
86
Coding
43
Agentic
24
Long context
70

4 tested benchmark families

Benchmark placements

Where Qwen3 Max Thinking places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
IFBenchReasoning#50/ 12771Artificial Analysis2026-09-01
GPQA DiamondReasoning#83/ 18586Artificial Analysis2026-09-01
AA-LCRLong context#88/ 17970Artificial Analysis2026-09-01
Humanity's Last ExamReasoning#88/ 18028Artificial Analysis2026-09-01
SciCodeCoding#96/ 18043Artificial Analysis2026-09-01
Artificial Analysis Intelligence IndexReasoning#131/ 18433Artificial Analysis2026-09-01
Terminal-BenchAgentic#164/ 18424Artificial Analysis2026-09-01
AA output speedPerformance#180/ 185Artificial Analysis2026-09-01
AA time to first tokenPerformance#180/ 185Artificial Analysis2026-09-01
Family context

The three highest-scoring models with pages in each capability family. Where Qwen3 Max Thinking shows up, it's highlighted.

Newest receipts

Benchmark rows added to the public ledger for Qwen3 Max Thinking in the last 120 days. Older rows live in the full table below.

9 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Reasoning86
Coding43
Agentic24
Long context70
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMetaMiniMaxxAIZhipuXiaomiOtherMeituanCohereNVIDIA

Full benchmark ledger

Every catalog benchmark for Qwen3 Max Thinking. Scores link to the original source; gaps mean no public row exists yet.

9 sourced rows

9 of 37 catalog benchmarks have a sourced row for Qwen3 Max Thinking.

Reasoning

Coding

Agentic

Long context

Performance

Sources