Models / Zhipu

GLM-5.2

GLM · Released 2026-06-13

Z.ai's June 2026 flagship for long-horizon coding. 1M context, 128K max output, Terminal-Bench 2.1 81.0, SWE-bench Pro 62.1. Highest-ranked open-source model on FrontierSWE, PostTrainBench, and SWE-Marathon. $1.40 / $4.40 per 1M tokens, MIT weights pending.

Long-horizon coding and autonomous agent model for project-scale engineering. The GLM pick when you need 1M context that stays usable across multi-day refactors, not just a large window.

#16 of 168 on AA Index · current snapshot
53 AA Index · 12 benchmark rows · 3 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite60
Context1M
Input / 1M$1.4
Output / 1M$4.4
Knowledge cutoffopen
Output speed69 t/s
TTFT (default API)1.36s
Statussolid

Family profile

Best published score in each covered benchmark family.

74/100 avg
Reasoning
90
Coding
51
Agentic
78
Long context
77

4 tested benchmark families

Editor's note
GLM-5.2 is the strongest open-source coding model we track. Terminal-Bench 2.1 at 81.0 lands within 4 points of Claude Opus 4.8 (85.0), and the FrontierSWE gap to Opus is 1%. We keep the tool at solid until independent benchmarks beyond Z.ai's own reports land in the ledger. MIT weights are announced but not yet shipped — verify before pinning self-hosted workflows.
Benchmark placements

Where GLM-5.2 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
Vending-Bench 2Agentic#4/ 10Andon Labs Vending-Bench 22026-08-04
DeepSWECoding#7/ 844DeepSWE leaderboard v1.12026-06-20
Artificial Analysis Intelligence IndexReasoning#26/ 18753Artificial Analysis2026-09-01
Humanity's Last ExamReasoning#29/ 18041Artificial Analysis2026-09-01
τ³-Bench BankingAgentic#29/ 10635Artificial Analysis2026-09-01
Terminal-BenchAgentic#30/ 18478Artificial Analysis2026-09-01
AA-LCRLong context#32/ 17977Artificial Analysis2026-09-01
IFBenchReasoning#35/ 12773Artificial Analysis2026-09-01
SciCodeCoding#39/ 18051Artificial Analysis2026-09-01
GPQA DiamondReasoning#49/ 18590Artificial Analysis2026-09-01
AA time to first tokenPerformance#50/ 186Artificial Analysis2026-09-01
AA output speedPerformance#56/ 187Artificial Analysis2026-09-01
Family context

The three highest-scoring models with pages in each capability family. Where GLM-5.2 shows up, it's highlighted.

Newest receipts

Benchmark rows added to the public ledger for GLM-5.2 in the last 120 days. Older rows live in the full table below.

12 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Reasoning90
Coding51
Agentic78
Long context77
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover any dot for name, score, and input price. Keyboard: tab through the top twelve, or use the ranking below.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMetaMiniMaxxAIZhipuXiaomiOtherMeituanCohereNVIDIA

Full benchmark ledger

Every catalog benchmark for GLM-5.2. Scores link to the original source; gaps mean no public row exists yet.

12 sourced rows

12 of 37 catalog benchmarks have a sourced row for GLM-5.2.

Reasoning

Coding

Agentic

Long context

Performance

Changelog
  1. Released GLM-5.2 with 1M context, 128K max output, two thinking modes, and long-horizon coding focus. MIT open weights announced for the following week.
Sources