Models / Alibaba

Qwen 3.7 Max

Qwen · Released 2026-05-20

Alibaba's May 2026 flagship open-weights model, matching frontier reasoning at open-weight price points. AA Intelligence Index 56.6 (#6 globally), Terminal-Bench 2.0 69.7%. 128K context, $0.40 / $1.20 per 1M tokens.

Default model for cost-sensitive deployments, multilingual applications (especially Chinese), and teams that need open weights with frontier-tier scores.

#15 of 21 on Vals Index · current snapshot
48 Vals Index · 14 benchmark rows · 9 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite73
Context128K
Input / 1M$0.40
Output / 1M$1.2
Knowledge cutoff2026-02-01
Statusflagship
KnowledgeReasoningMathCodingAgenticLong contextTool useHuman preference
Knowledge
89
Reasoning
74
Math
82
Coding
93
Agentic
70
Long context
48
Tool use
78
Human preference
82

77average across 8 tested families

Editor's note
Qwen 3.7 Max is the strongest open-weights model from any provider outside the US. If you need to run a model on your own metal and Chinese-language support matters, this is the one to benchmark first.
Benchmark placements

Where Qwen 3.7 Max places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
LMArena (Chatbot Arena)Human preference#7/ 1382LMArena leaderboard2026-05-25
MCP AtlasTool use#7/ 1072MCP Atlas leaderboard2026-06-09
BFCLTool use#10/ 2078Berkeley Function Calling Leaderboard V42026-04-15
HumanEvalCoding#10/ 1593CodeSOTA HumanEval leaderboard2026-05-20
MMLUKnowledge#10/ 2289Alibaba Qwen 3.7 Max2026-05-20
SWE-bench VerifiedCoding#14/ 2379SWE-bench Verified leaderboard2026-06-09
Vals IndexAgentic#14/ 2048Vals AI — Vals Index2026-06-04
LiveBenchReasoning#15/ 2162LiveBench leaderboard2026-04-15
MATHMath#15/ 2082Alibaba Qwen 3.7 Max2026-05-20
LongBench v2Long context#16/ 1948LongBench v2 leaderboard2026-04-01
MMLU-ProKnowledge#16/ 2175MMLU-Pro HuggingFace leaderboard2026-05-20
Terminal-BenchAgentic#22/ 6970Terminal-Bench 2.0 leaderboard2026-05-20
GPQA DiamondReasoning#60/ 6974Artificial Analysis GPQA Diamond evaluation2026-06-09
Family context

The three highest-scoring carded models in each capability family. Where Qwen 3.7 Max shows up, it's highlighted.

KnowledgeQwen 3.7 Max · 89
  1. 1
    91
  2. 2
    91
  3. 3
    91
ReasoningQwen 3.7 Max · 74
  1. 1
    95
  2. 2
    94
  3. 3
    94
MathQwen 3.7 Max · 82
  1. 1
    99
  2. 2
    97
  3. 3
    96
CodingQwen 3.7 Max · 93
  1. 1
    96
  2. 2
    96
  3. 3
    96
AgenticQwen 3.7 Max · 70
  1. 1
    100
  2. 2
    100
  3. 3
    100
Long contextQwen 3.7 Max · 48
  1. 1
    75
  2. 2
    75
  3. 3
    74
Tool useQwen 3.7 Max · 78
  1. 1
    89
  2. 2
    87
  3. 3
    87
Human preferenceQwen 3.7 Max · 82
  1. 1
    95
  2. 2
    92
  3. 3
    91
Newest receipts

Benchmark rows added to the public ledger for Qwen 3.7 Max in the last 120 days. Older rows live in the full table below.

14 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Knowledge89
Reasoning74
Math82
Coding93
Agentic70
Long context48
Tool use78
Human preference82
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMiniMaxxAIMetaZhipuXiaomiMeituanCohere

Full benchmark ledger

Every catalog benchmark for Qwen 3.7 Max. Scores link to the original source; gaps mean no public row exists yet.

14 sourced rows

13 of 32 catalog benchmarks have a sourced row for Qwen 3.7 Max.

Knowledge

Reasoning

Math

Coding

Agentic

Long context

Tool use

Human preference

Changelog
  1. Released Qwen 3.7 Max — AA Index 56.6, Terminal-Bench 2.0 69.7%.
  2. Qwen 3.7 Plus retired for general availability.
Sources