Models / Alibaba

Qwen3.8 2.4T A95B

Qwen · Released 2026-08-12

Alibaba's open-weight Qwen-Max-class MoE (12 Aug 2026): 2.4T total / 95B active. Artificial Analysis Intelligence Index 57.7. AA list price $2 / $6 per 1M tokens.

Self-hostable Qwen3.8 checkpoint for long-context and agentic work. Hosted Qwen3.8 Max adds vision and a non-thinking mode. Prefer sourced ledger rows over vendor eval tables.

#7 of 166 on AA Index · current snapshot
58 AA Index · 9 benchmark rows · 1 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite64
Context262K
Input / 1M$2.0
Output / 1M$6.0
Knowledge cutoffunpublished
Output speed40 t/s
TTFT (default API)1.76s
Statussolid

Family profile

Best published score in each covered benchmark family.

76/100 avg
Reasoning
94
Coding
52
Agentic
82
Long context
75

4 tested benchmark families

Editor's note
Recorded from the 2026-08-15 AA snapshot (slug qwen3-8-2-4t-a95b, index 57.7) plus the Hugging Face model card. This is the open-weight Max-class checkpoint, not the hosted Qwen3.8 Max API product. Editorial depth still needs a desk pass.
Benchmark placements

Where Qwen3.8 2.4T A95B places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
τ³-Bench BankingAgentic#4/ 10649Artificial Analysis2026-09-01
GPQA DiamondReasoning#9/ 18594Artificial Analysis2026-09-01
Artificial Analysis Intelligence IndexReasoning#12/ 18458Artificial Analysis2026-09-01
Terminal-BenchAgentic#16/ 18482Artificial Analysis2026-09-01
Humanity's Last ExamReasoning#24/ 18042Artificial Analysis2026-09-01
AA time to first tokenPerformance#33/ 185Artificial Analysis2026-09-01
SciCodeCoding#34/ 18052Artificial Analysis2026-09-01
AA-LCRLong context#47/ 17975Artificial Analysis2026-09-01
AA output speedPerformance#78/ 185Artificial Analysis2026-09-01
Family context

The three highest-scoring models with pages in each capability family. Where Qwen3.8 2.4T A95B shows up, it's highlighted.

Newest receipts

Benchmark rows added to the public ledger for Qwen3.8 2.4T A95B in the last 120 days. Older rows live in the full table below.

9 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Reasoning94
Coding52
Agentic82
Long context75
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMetaMiniMaxxAIZhipuXiaomiOtherMeituanCohereNVIDIA

Full benchmark ledger

Every catalog benchmark for Qwen3.8 2.4T A95B. Scores link to the original source; gaps mean no public row exists yet.

9 sourced rows

9 of 37 catalog benchmarks have a sourced row for Qwen3.8 2.4T A95B.

Reasoning

Coding

Agentic

Long context

Performance

Changelog
  1. Open weights published. AA Intelligence Index 57.7; hosted list $2 / $6 per 1M tokens.
Sources