Models / Zhipu

GLM-5.1

GLM · Released 2026-04-07

Z.ai's April 2026 flagship for long-horizon agentic engineering. Official docs list text input/output, 200K context, 128K maximum output, function calling, structured output, context caching, MCP, and open weights under the MIT license.

Long-horizon coding and autonomous agent model for Claude Code/OpenClaw-style workflows when you want open weights from the GLM line.

#59 of 168 on AA Index · current snapshot
41 AA Index · 11 benchmark rows · 2 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite51
Context200K
Input / 1M
Output / 1M
Knowledge cutoffopen
Statussolid

Family profile

Best published score in each covered benchmark family.

59/100 avg
Reasoning
87
Math
33
Coding
44
Agentic
62
Long context
68

5 tested benchmark families

Editor's note
GLM-5.1 is the GLM tool to watch for long autonomous coding runs. Z.ai's own benchmark claims are strong, but we keep the tool at solid until independent score rows are imported into VerdictPal's ledger.
Benchmark placements

Where GLM-5.1 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
FrontierMathMath#14/ 1533Epoch AI FrontierMath Tiers 1-3 (v2) CSV2026-08-15
IFBenchReasoning#18/ 12776Artificial Analysis2026-09-01
Terminal-BenchAgentic#63/ 18462Artificial Analysis2026-09-01
Artificial Analysis Intelligence IndexReasoning#68/ 18741Artificial Analysis2026-09-01
Humanity's Last ExamReasoning#74/ 18030Artificial Analysis2026-09-01
GPQA DiamondReasoning#75/ 18587Artificial Analysis2026-09-01
τ³-Bench BankingAgentic#77/ 10614Artificial Analysis2026-09-01
SciCodeCoding#88/ 18044Artificial Analysis2026-09-01
AA-LCRLong context#106/ 17968Artificial Analysis2026-09-01
AA time to first tokenPerformance#122/ 186Artificial Analysis2026-09-01
AA output speedPerformance#123/ 187Artificial Analysis2026-09-01
Family context

The three highest-scoring models with pages in each capability family. Where GLM-5.1 shows up, it's highlighted.

Newest receipts

Benchmark rows added to the public ledger for GLM-5.1 in the last 120 days. Older rows live in the full table below.

11 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Reasoning87
Math33
Coding44
Agentic62
Long context68
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover any dot for name, score, and input price. Keyboard: tab through the top twelve, or use the ranking below.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMetaMiniMaxxAIZhipuXiaomiOtherMeituanCohereNVIDIA

Full benchmark ledger

Every catalog benchmark for GLM-5.1. Scores link to the original source; gaps mean no public row exists yet.

11 sourced rows

11 of 37 catalog benchmarks have a sourced row for GLM-5.1.

Reasoning

Math

Coding

Agentic

Long context

Performance

Changelog
  1. Released GLM-5.1 with 200K context, 128K max output, long-horizon coding focus, and MIT-licensed open weights.
Sources