Models / Zhipu

GLM-5.2

GLM · Released 2026-06-13

Z.ai's June 2026 flagship for long-horizon coding. 1M context, 128K max output, Terminal-Bench 2.1 81.0, SWE-bench Pro 62.1. Highest-ranked open-source model on FrontierSWE, PostTrainBench, and SWE-Marathon. $1.40 / $4.40 per 1M tokens, MIT weights pending.

Long-horizon coding and autonomous agent model for project-scale engineering. The GLM pick when you need 1M context that stays usable across multi-day refactors, not just a large window.

#8 of 51 on AA Index · current snapshot
51 AA Index · 12 benchmark rows · 3 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite58
Context1M
Input / 1M$1.4
Output / 1M$4.4
Knowledge cutoffopen
Output speed196 t/s
TTFT (default API)0.84s
Statussolid
ReasoningCodingAgenticLong context
Reasoning
90
Coding
51
Agentic
78
Long context
71

73average across 4 tested families

Editor's note
GLM-5.2 is the strongest open-source coding model we track. Terminal-Bench 2.1 at 81.0 lands within 4 points of Claude Opus 4.8 (85.0), and the FrontierSWE gap to Opus is 1%. We keep the card at solid until independent benchmarks beyond Z.ai's own reports land in the ledger. MIT weights are announced but not yet shipped — verify before pinning self-hosted workflows.
Benchmark placements

Where GLM-5.2 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
Vending-Bench 2Agentic#2/ 5Andon Labs Vending-Bench 22026-06-01
DeepSWECoding#5/ 644DeepSWE leaderboard v1.12026-06-20
AA output speedPerformance#12/ 69Artificial Analysis2026-07-21
Artificial Analysis Intelligence IndexReasoning#13/ 6751Artificial Analysis2026-07-21
Humanity's Last ExamReasoning#14/ 6540Artificial Analysis2026-07-21
Terminal-BenchAgentic#14/ 6978Artificial Analysis2026-07-21
AA-LCRLong context#15/ 6471Artificial Analysis2026-07-21
IFBenchReasoning#24/ 5773Artificial Analysis2026-07-21
SciCodeCoding#25/ 6551Artificial Analysis2026-07-21
GPQA DiamondReasoning#27/ 6990Artificial Analysis2026-07-21
τ³-Bench BankingAgentic#32/ 6427Artificial Analysis2026-07-21
AA time to first tokenPerformance#49/ 69Artificial Analysis2026-07-21
Family context

The three highest-scoring carded models in each capability family. Where GLM-5.2 shows up, it's highlighted.

Newest receipts

Benchmark rows added to the public ledger for GLM-5.2 in the last 120 days. Older rows live in the full table below.

12 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Reasoning90
Coding51
Agentic78
Long context71
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMiniMaxxAIMetaZhipuXiaomiMeituanCohere

Full benchmark ledger

Every catalog benchmark for GLM-5.2. Scores link to the original source; gaps mean no public row exists yet.

12 sourced rows

12 of 32 catalog benchmarks have a sourced row for GLM-5.2.

Reasoning

Coding

Agentic

Long context

Performance

Changelog
  1. Released GLM-5.2 with 1M context, 128K max output, two thinking modes, and long-horizon coding focus. MIT open weights announced for the following week.
Sources