GLM-5: from Vibe Coding to Agentic Engineering | Summary
17 Feb 2026 | Paper Review AI Coding Agents Reinforcement Learning Sparse Attention Environment ScalingContents
- Summary
- 1 Introduction
- 2 Pre-Training
- 3 Post-Training
- 4 Agentic Engineering
- 5 Adapting GLM-5 to Chinese Chip Infrastructure
- 6 Evaluation
- 7 Conclusion
- 8 Easter Eggs
- 9 Contribution
- Appendix
- Brief Thoughts
This article explains the key points of GLM-5: from Vibe Coding to Agentic Engineering.
- 2026-02-17 (arXiv)
- GLM-5-Team, :, Zeng, Aohan, Lv, Xin, Hou, Zhenyu, Du, Zhengxiao, Zheng, Qinkai, Chen, Bin, Yin, Da, Ge, Chendi, et al.
- Zhipu AI, Tsinghua University
- Paper
- Github
- Project Page
Summary
- GLM-5: from Vibe Coding to Agentic Engineering presents an open-weight foundation model for agents that plan, implement, and revise software over extended interactions. Its 744B-parameter MoE architecture activates 40B parameters, and its base-model training uses a reported 28.5 trillion tokens with context extension to 200K.
- The training pipeline combines DeepSeek Sparse Attention (DSA), sequential reasoning, agentic, and general reinforcement learning, and final on-policy cross-stage distillation. Asynchronous agent training separates rollout generation from optimization, while exact token preservation, importance sampling, stale-trajectory filtering, and executable environments address policy mismatch and unreliable feedback.
- GLM-5 scores 77.8 on SWE-bench Verified, 73.3 on SWE-bench Multilingual, and 75.9 on BrowseComp with context management. Internal engineering evaluations improve over GLM-4.7, but complete frontend task success and chained development remain below Claude Opus 4.5. Modified benchmark conditions, proprietary judges, and limited component-level ablations restrict conclusions about general engineering reliability and system efficiency.
1 Introduction
The introduction identifies computational cost and adaptation to complex software workflows as bottlenecks beyond model scaling. GLM-5 targets unified agentic, reasoning, and coding (ARC) capabilities, emphasizing autonomous planning and iteration rather than isolated code generation.
- The report records an Artificial Analysis Intelligence Index v4.0 score of 50, compared with 42 for GLM-4.7, and first place among open models in the displayed LMArena Text Arena and Code Arena snapshots.
- Vending-Bench 2 evaluates a simulated business over a year; GLM-5 finishes with $4,432, below Claude Opus 4.5 at $4,967 and Gemini 3 Pro at $5,478.
- The stated contributions cover sparse attention, asynchronous RL infrastructure and algorithms, and deployment adaptation across seven Chinese chip platforms. Code and models are linked at https://github.com/zai-org/GLM-5.
The eight panels show competitive ARC performance rather than uniform dominance. GLM-5 scores 77.8 on SWE-bench Verified versus 80.9 for Claude Opus 4.5, while its displayed BrowseComp score is 75.9. The image and caption identify DeepSeek-V3.2 as a comparator despite the introduction’s reference to GLM-4.7; some BrowseComp comparator values also differ from Table 7.
The leaderboard snapshot places GLM-5 at an Intelligence Index score of 50, up from GLM-4.7 at 42. It supplies an external aggregate evaluation signal, but the ten-task index does not measure complete software-workflow reliability.
The snapshots identify GLM-5 as the highest-ranked open model while showing overall positions of #11 in Text Arena and #6 in Code Arena. Open-model rank should therefore not be interpreted as overall first place.
The business-simulation trajectories show GLM-5 reaching a final balance of $4,432. The engineering panels show gains over GLM-4.7 but a remaining chained-task gap to Claude Opus 4.5, distinguishing long-horizon progress from parity on sequential development.
2 Pre-Training
Base-model development separates general language and code pre-training from mid-training for long-context and agentic capabilities. The pipeline depicts 18T general pre-training tokens and 9T code and reasoning tokens at 4K context, followed by progressively longer-context training and sparse-attention adaptation. The report summarizes the overall base-model budget as 28.5 trillion tokens; this rounded total is not an exact sum of the displayed stage budgets.
The diagram places 4K pre-training before 32K, 128K, and 200K mid-training and 20B-token sparse adaptation. Post-training moves from SFT through three RL stages, with earlier-stage logits feeding final cross-stage distillation.
2.1 Architecture
GLM-5 scales to 256 experts and a reported 80 layers, with the reduced depth intended to limit expert-parallel communication. Multi-latent Attention (MLA) reduces KV-cache storage, but its interaction with the Muon optimizer motivates head-specific orthogonalization, called Muon Split.
- Muon Split orthogonalizes separate query, key, and value up-projection matrices for individual heads instead of orthogonalizing each combined multi-head matrix. Table 1 shows gains on most listed benchmarks over ordinary MLA, although GSM8K declines and the results do not uniformly match GQA-8.
- MLA-256 increases the head dimension from 192 to 256 and reduces the number of heads by 1/3, retaining training computation and parameter count while reducing decoding computation.
- Three training-time Multi-token Prediction (MTP) layers share parameters; Table 10 lists one MTP layer under this architecture. With four speculative steps on a private prompt set, GLM-5 achieves an acceptance length of 2.76 versus 2.55 for DeepSeek-V3.2.
Muon Split raises MLA MMLU from 61.5 to 62.5 and HumanEval from 33.5 to 36.7, but GSM8K falls from 46.2 to 45.0. The mixed results support the optimizer adaptation without establishing a uniform improvement on every task.
With four speculative steps on the authors’ private prompt set, GLM-5 accepts an average length of 2.76 compared with 2.55 for DeepSeek-V3.2. This supports higher acceptance on that workload, but does not itself measure end-to-end latency.
2.1.1 Continued Pre-Training with DeepSeek Sparse Attention (DSA) converts the dense MLA model after mid-training through indexer warmup and sparse adaptation. Warmup uses 1000 steps, 14 sequences of 202,752 tokens per step, and a maximum learning rate of 5e-3; adaptation then uses 20B tokens.
- DSA selects content-relevant tokens rather than imposing a fixed sliding window. The section reports roughly 1.5–2× lower attention computation for long sequences, but does not provide a matched GLM-5 end-to-end cost measurement.
- Table 3 shows comparable but non-identical long-context results: DSA improves SQuAD-128k from 79.7 to 86.0 but reduces HotpotQA-128k from 66.3 to 63.0.
- Fine-tuning MLA and DSA on the same SFT data produces nearly overlapping loss curves, and the authors report tied evaluation performance. These observations support successful adaptation on the measured workloads, not a universal guarantee of lossless sparsification.
DSA preserves MQ-NIAH-128k at 100.0, improves MV-NIAH-128k and SQuAD-128k, and lowers HotpotQA-128k from 66.3 to 63.0. The results show broadly comparable long-context behavior with task-dependent differences.
2.1.2 Ablation Study of Efficient Attention Variants compares fixed and searched Sliding Window Attention (SWA), Gated DeltaNet (GDN), and SimpleGDN on GLM-9B. The efficient-attention variants undergo 190B-token continual training at 64K context with a 1:1 ratio of efficient-attention and full-attention layers; the separate DSA experiment uses GLM-4.7-Flash.
- SWA layer search uses beam size 8, changes two layers per step, and evaluates candidates on RULER at 16K. Its resulting pattern is SFSSFFSSSFFFFSSFSFFFFFFSFSFSSFSSFSFSSFSSS, where S denotes SWA and F denotes full attention.
- Without additional training, a 4096-token window and searched placement give 53.95 on RULER at 128K, compared with 6.51 for fixed interleaving and 75.28 for full attention.
- After continual training, searched SWA reaches 69.59 on RULER at 128K; SimpleGDN reaches 67.03 and improves HELMET-ICL to 81.84 versus 77.36 for the baseline. SimpleGDN removes Conv1d and explicit gating to reuse pretrained projections without adding parameters.
- On GLM-4.7-Flash, indexer-only DSA warmup reduces RULER at 128K from 79.21 to 71.35; joint training on 150B tokens recovers it to 78.86. Different backbones and training budgets prevent a controlled ranking of DSA against the GLM-9B alternatives.
At 128K, searched SWA placement achieves 53.95 compared with 6.51 for fixed interleaving, despite the same 1:1 attention-layer ratio and 4096-token window. The full-attention baseline remains higher at 75.28, showing that placement mitigates but does not remove retrieval degradation.
Following 190B-token continual training, searched SWA and SimpleGDN perform better than naive interleaving on the listed long-context tests. SimpleGDN exceeds the baseline on HELMET-ICL but remains lower on RULER, MRCR, and RepoQA, demonstrating task-dependent accuracy trade-offs.
2.2 Pre-training Data
The pre-training corpus expands web, code, and math and science coverage through domain-specific quality filtering. A sentence-embedding-based DCLM classifier supplements existing web selection, while a World Knowledge classifier extracts long-tail information from otherwise medium- or low-quality documents.
- Code collection adds refreshed hosting-platform snapshots and code-containing webpages, increasing fuzzily deduplicated unique tokens by 28%. Metadata alignment, language identification, and dedicated low-resource-language classifiers improve corpus integrity and sampling.
- Math and science selection combines improved webpage extraction and PDF parsing with LLM educational-quality scoring. Long documents use chunk-and-aggregate scoring, and this corpus explicitly filters synthetic, AI-generated, and template-based material.
2.3 Mid-Training
Mid-training extends context through 32K with 1T tokens, 128K with 500B tokens, and 200K with 50B tokens. Later stages upsample long documents and synthetic agent trajectories to support multi-file engineering and extended interactions.
- Repository sequences concatenate code, commit diffs, issues, pull requests, and relevant files. Approximately 10 million issue–PR pairs are collected, and the filtered issue–PR portion contains approximately 160B unique tokens.
- Natural long-context material undergoes perplexity, deduplication, and length filtering. Synthetic sequences use interleaved packing of similar texts to construct long-range dependencies.
- A small proportion of MRCR-like data is introduced at 200K. The authors report that this additional stage improves performance even within the earlier 128K window, without a numerical ablation in this section.
2.4 Training Infrastructure
Training infrastructure addresses memory imbalance and communication overhead through MTP placement, gradient sharding, optimizer communication changes, activation offloading, and chunked output projection. Parallelism optimizations defer weight-gradient computation and rebalance long-sequence workloads; INT4 quantization-aware training is applied during SFT.
- The MTP output shares placement with the main output on the final pipeline stage, while its embedding and transformer components occupy the preceding stage.
- Pipeline ZeRO2 stores a 1/dp fraction of each stage’s gradients and retains only two full accumulation buffers through double buffering. Muon communication gathers only parameter shards owned by each rank and overlaps communication with local computation.
- Layer-granular host offloading and recomputation reduce resident activation memory. Sequence-chunked projection and loss computation release intermediate activations before processing the next chunk.
- Sequence reordering, variable-size context-parallel groups, and hierarchical all-to-all address long-context imbalance. A shared quantization kernel provides bitwise-identical quantization behavior between training and inference.
3 Post-Training
Post-training begins with multi-task Supervised Fine-Tuning (SFT), then proceeds through Reasoning RL, Agentic RL, and General RL. Final on-policy cross-stage distillation is intended to recover capabilities degraded by sequential optimization for different objectives.
3.1 Supervised Fine-Tuning
SFT covers General Chat, Reasoning, and Coding & Agent data, with a maximum context length of 202,752 tokens. Its updated chat template supports Interleaved Thinking before responses and tool calls, Preserved Thinking across coding-agent turns, and per-turn reasoning control through Turn-level Thinking.
- General Chat data emphasizes concise responses and broader multilingual roleplay coverage, with automatic and human filtering. Reasoning examples include verifiable logical problems generated through rejection sampling and mathematical or scientific problems filtered for difficulty relative to GLM-4.7.
- Coding and agent trajectories come from execution environments, expert RL, and rejection sampling, emphasizing realistic long-horizon tasks.
- Erroneous trajectory segments remain available as context but are masked from the training loss, allowing subsequent recovery actions to provide supervision without reinforcing the mistakes.
The diagram distinguishes reasoning inserted between tool actions from reasoning retained across user turns. Preserved Thinking keeps earlier reasoning blocks in the next turn’s context, whereas the alternative retains tool interactions and answers without those blocks.
3.2 Reasoning RL
Reasoning RL builds on GRPO and IcePop, distinguishing training-policy probabilities from inference-policy probabilities and removing KL regularization. It uses β = 2, ϵlow = 0.2, ϵhigh = 0.28, group size 32, and batch size 32 in an entirely on-policy setting.
- IcePop retains the training–inference probability ratio as a weight when it lies within [1/β, β] and otherwise assigns zero. The GRPO objective separately uses clipped current-to-old training-policy ratios and group-normalized outcome advantages.
- For DSA, storing every selected index would be expensive because the indexer selects k = 2048 entries per token. The authors instead use torch.topk, which they report as deterministic in their implementation, and freeze indexer parameters by default during RL.
- The tested non-deterministic CUDA and TileLang top-k implementations reportedly cause rapid performance deterioration and entropy collapse, making consistent sparse-token selection a practical stability requirement.
- The training mixture is roughly balanced across mathematics, science, code, and tool-integrated reasoning. Domain-specific judges or execution systems provide binary outcome rewards.
3.3 Agentic RL
Agentic RL jointly trains coding and search agents using a fully asynchronous framework with a central Multi-Task Rollout Orchestrator. A Token-in-Token-out (TITO) gateway preserves sampled token identities, and Direct Double-sided Importance Sampling masks tokens with excessive policy divergence without retaining historical policy checkpoints.
- DP-aware routing preserves KV-cache locality for long-context MoE inference.
- The training tasks draw on real-world SWE environments, terminal environments, and difficult multi-hop search questions; their construction and asynchronous stability mechanisms are detailed in Section 4.
3.4 General RL
General RL separates foundational correctness, emotional intelligence, and task-specific quality. Correctness targets instruction failures, logical inconsistency, factual errors, hallucinations, and disfluency; emotional intelligence targets empathetic, natural responses; task-specific rewards refine writing, question answering, roleplay, translation, and related tasks.
- The hybrid reward system combines deterministic rules, outcome reward models (ORMs), and generative reward models (GRMs). The report contrasts rule precision, ORM efficiency and susceptibility to hacking, and GRM robustness and higher variance.
- Expert human-authored responses serve as stylistic anchors intended to reduce verbose or formulaic model-generated patterns.
- The section describes reward design without isolating the numerical contribution of each reward type or the human-response component.
3.5 On-Policy Cross-Stage Distillation
On-Policy Cross-Stage Distillation uses final checkpoints from preceding stages as teachers and mixes prompts from their corresponding training sets. It replaces outcome-based advantages with a stop-gradient log-probability ratio between the teacher’s inference policy and the current student’s training policy.
- The student generates the trajectories on which teacher probabilities are evaluated, keeping supervision aligned with the student’s current behavior.
- The GRPO group size becomes 1 and batch size becomes 1024 because teacher-derived advantages do not require multiple responses to estimate a group baseline.
- Teacher logits currently come from an inference engine. Migration to a training-engine backend and uniform MQA-mode MLA inference is described as future work, not an implemented result.
3.6 RL Training Infrastructure: The slime Framework
The slime framework supplies a common post-training stack for customized rollouts, HTTP-based task execution, latency-oriented inference, and fault tolerance. The report emphasizes completion time of the slowest trajectories rather than aggregate serving throughput alone.
- Task-specific rollout code can implement multi-turn interactions, tool calls, environmental feedback, and verifier-guided branching without changing the optimization backend. HTTP APIs also let external agent frameworks use the same serving layer.
- A multi-node deployment example uses EP64 and DP64 over 8 nodes to provide distributed KV-cache capacity. FP8 rollouts and MTP target decoding latency, particularly for small-batch stragglers.
- Prefill–Decode disaggregation separates long-prefix prefills from ongoing decoding to reduce interference in multi-turn workloads.
- Heartbeat monitoring terminates and deregisters unhealthy servers so retries reach healthy servers. The section supplies design details but no measured throughput or tail-latency comparison against a synchronous baseline.
4 Agentic Engineering
The report defines agentic engineering as agents planning, implementing, and iterating, rather than humans repeatedly prompting a model to write code. Its training approach couples asynchronous RL with environment-building pipelines that supply executable feedback for software engineering, terminal use, search, and slide generation.
4.1 Asynchronous RL for Agentic Tasks
Asynchronous agent training places inference and training engines on separate GPUs: inference continuously collects trajectories, and optimization starts when a predefined collection threshold is reached. Rollout weights synchronize every K gradient updates, with optimizer resets after inference-engine weight updates; a central orchestrator supports over 1k concurrent rollouts across independent task microservices.
- Only model-generated tokens enter the optimization loss; environment feedback remains context. TITO records exact token IDs and metadata to avoid token-boundary, normalization, and truncation mismatches caused by re-tokenizing finalized text.
- Direct importance sampling uses the current-policy probability divided by the recorded rollout probability. Its calibration function retains this ratio only when 1 − ϵℓ < x < 1 + ϵh and otherwise masks the token from gradient computation.
- A trajectory is discarded when its oldest rollout version satisfies w′ − w0 > τ. Failures caused by environment collapse are excluded; incomplete groups repeat valid samples only when more than half the original group remains, otherwise the whole group is dropped.
- Consistent hashing assigns each rollout to a stable DP rank, with lightweight rebalancing to preserve prefix-cache locality. These mechanisms address policy lag and heterogeneous rollout durations, but their individual gains are not quantified.
4.2 Environment Scaling for Agents
4.2.1 Software Engineering (SWE) Environments and 4.2.2 Terminal Environments construct tasks with executable verification rather than relying only on textual judgments. The SWE pipeline builds over 10k environments across thousands of repositories and 9 languages: Python, Java, Go, C, CPP, JavaScript, TypeScript, PHP, and Ruby.
- Filtered issue–PR pairs cover bug fixing, feature implementation, refactoring, and other changes. A RepoLaunch-based setup pipeline generates environments and test commands, with language-aware log parsing for Fail-to-Pass (F2P) and Pass-to-Pass (P2P) cases.
- Seed-based terminal synthesis proceeds through drafting, concrete Harbor-format implementation, and rubric-guided refinement. It produces thousands of tasks and reports Docker construction accuracy exceeding 90%.
- A second terminal pipeline selects technical webpages, samples across topics and difficulty, and lets a construction agent revise its task until Harbor validation passes. These checks establish construction validity but do not independently demonstrate complete resistance to shortcuts.
4.2.3 Search Tasks synthesizes multi-hop questions from a Web Knowledge Graph built from over two million deduplicated webpages. 4.2.4 Inference with Context Management for Search Agents shows that managing accumulated tool observations remains important despite the model’s extended context window.
- Question filtering removes items solved in any of eight tool-free attempts or by a basic agent in a few steps. Bidirectional verification rejects non-unique answers, inconsistent evidence, and incorrect labels.
- Search evaluation uses the official OpenAI judging prompt and o3-mini after the authors observe sensitivity to judge choice. The tools are search, open, find, and python.
- Keep-recent-k replaces older observations with the exact text “Tool result is omitted to save tokens.” With k = 5, BrowseComp improves from 55.3% to 62.0%.
- Hierarchical Context Management combines observation folding with Discard-all when total context exceeds T = 32k, reaching 75.9. Section 6.1.3 and Appendix B.2 instead describe the reported context-management condition as discard-all, leaving the precise benchmark configuration unresolved.
BrowseComp accuracy depends on both search strategy and executed-step budget. GLM-5’s Hierarchical Context Management improves over the plotted GLM-4.7 baselines; the separate Pass@K curve represents a different evaluation condition and should not be read as single-trajectory accuracy.
4.2.5 Slide Generation trains an HTML-based slide expert through SFT, RL, rejection sampling, and masking-based refinement. Its reward hierarchy evaluates static markup, runtime DOM geometry, and rendered visual features, with renderer revisions addressing content truncation and spacing manipulation.
- Dynamic sampling reduces the share of structurally trivial pages. Best-of-N rejection sampling selects high-quality candidates, while masking preserves useful pages from otherwise partially defective trajectories.
- Strict 16:9 compliance increases from 40% to 92%. Against GLM-4.5, human evaluation reports win rates of 60% for content, 57.5% for layout, 65% for aesthetics, and 67.5% overall.
- The section documents concrete reward-hacking examples, but does not provide the human-evaluation sample size or uncertainty for the win rates.
One example changes a 1280×1015 slide into 1280×720 by hiding overlong content; another expands a 1280×493 layout to 1280×720 through added space. Both meet superficial size criteria without necessarily improving content or composition, motivating renderer-grounded reward checks.
5 Adapting GLM-5 to Chinese Chip Infrastructure
GLM-5 is adapted to Huawei Ascend, Moore Threads, Hygon, Cambricon, Kunlunxin, MetaX, and Enflame, with Ascend Atlas serving as the detailed case study. The deployment strategy combines mixed-precision W4A8 quantization, fused kernels, and inference-engine scheduling.
- On an Atlas 800T A3 machine, msModelSlim applies W8A8 to standard Attention and MLP blocks and W4A8 to MoE experts, with QuaRot and Flex_AWQ_SSZ supporting low-bit stability. This section calls the model 750B, whereas the architecture sections report 744B under their stated parameter-counting convention; the report does not explicitly reconcile the figures.
- Lightning Indexer fuses scoring, ReLU, and TopK; Sparse Flash Attention combines token selection and sparse computation; MLAPO fuses 13 preprocessing operators.
- vLLM-Ascend and SGLang optimizations include asynchronous scheduling, RadixCache and Prefix Cache, attention DP with expert EP, FlashComm, and MTP.
- The report claims a 50% deployment-cost reduction in long-sequence scenarios and performance comparable to “dual-GPU international clusters.” It supplies no workload, sufficiently specified hardware baseline, or cost breakdown for independently interpreting that comparison.
6 Evaluation
Evaluation combines public ARC benchmarks, the internal CC-Bench-V2 engineering suite, and five general-ability domains derived from deployment usage. Public comparisons include GLM-4.7, DeepSeek-V3.2, Kimi K2.5, Claude Opus 4.5, Gemini 3 Pro, and GPT-5.2 (xhigh).
6.1 Evaluation of ARC Benchmarks
6.1.1 Evaluation of Reasoning and General Benchmarks reports 30.5 on text-only HLE, 50.4 on HLE with tools, 97.9 on HMMT Feb. 2025, 96.9 on HMMT Nov. 2025, and 64.5 on LongBench v2. The LongBench v2 result is second to Gemini 3 Pro at 68.2, while Kimi K2.5 exceeds GLM-5 on both HLE conditions.
- Reasoning evaluations use a maximum generation length of 131,072 tokens. Appendix B.2 specifies a 202,752-token maximum context for HLE with tools; GPT-5.2 (medium) judges HLE.
- GLM-5 improves over GLM-4.7 on HLE and LongBench v2, but AIME 2026 I is slightly lower at 92.7 versus 92.9. Its IMO-AnswerBench and GPQA-Diamond scores are 82.5 and 86.0, respectively.
- Proprietary-model HLE-with-tools values marked with * use the full HLE set, whereas GLM-5 uses the text-only subset. These values are not matched-subset comparisons.
6.1.2 Evaluation of Coding Benchmarks reports 77.8 on SWE-bench Verified, 73.3 on SWE-bench Multilingual, 56.2 on the original Terminal-Bench 2.0 under both tested harnesses, and 43.2 on CyberGym. These results improve over GLM-4.7, while Claude Opus 4.5 remains higher on each of these original benchmark conditions.
- SWE evaluations use OpenHands with a tailored GLM-5 instruction prompt. Terminal-Bench uses Terminus-2 and Claude Code, with equal GLM-5 scores under the two reported original-set conditions.
- On a verified Terminal-Bench 2.0 version with ambiguous instructions corrected, GLM-5 scores 60.7 with Terminus-2 and 61.1 with Claude Code. These modified-set scores are not interchangeable with other models’ original-set results.
- The verified dataset is linked at https://huggingface.co/datasets/zai-org/terminal-bench-2-verified.
6.1.3 Evaluation of Agentic Abilities reports 62.0 on BrowseComp without full context resets, 75.9 with context management, 72.7 on BrowseComp-ZH, 89.7 on τ²-Bench, 67.8 on MCP-Atlas, and 39.2 on Tool-Decathlon. GLM-5 leads the reported BrowseComp comparisons, but does not uniformly lead tool-use or long-horizon benchmarks.
- MCP-Atlas compares models on the same 500-task public subset with a 10-minute timeout, increased from 4 minutes, and Gemini 3 Pro as judge.
- τ²-Bench includes Retail and Telecom simulator-prompt adjustments and Airline domain fixes, making simulator configuration part of the reported evaluation condition.
- Vending-Bench 2 yields $4,432, and GDPval-AA Elo is 1,409 as recorded on 15th Feb., 2026, compared with 1,400 for Claude Opus 4.5 and 1,462 for GPT-5.2 (xhigh).
The benchmark comparison shows gains over GLM-4.7, particularly in terminal use, CyberGym, and several tool-based tasks, alongside a small AIME 2026 I decline. Footnotes distinguish full-set HLE scores, verified Terminal-Bench scores, and date-specific GDPval-AA Elo, so not every entry is a matched-condition comparison.
6.2 Evaluation of Real-world Agentic Engineering Experience
CC-Bench-V2 measures frontend, backend, and long-horizon engineering through automated execution and Agent-as-a-Judge. Frontend projects undergo build verification followed by interactive testing with Claude Code, Claude Sonnet 4.5, and Playwright MCP. Scoring is automated, but task construction, unit tests, checklist review, and judge validation involve human experts.
- Build Success Rate (BSR) measures initialization and execution, Instance Success Rate (ISR) requires all specifications to pass, and Check-item Success Rate (CSR) measures individual requirement completion.
- Judge–human agreement is 94% on 130 sampled check-items. Across 8 frontier models, automated and human rankings have a reported Spearman correlation of 85.7%, with item-level disagreements concentrated on subjective visual quality.
- GLM-5 reports overall BSR of 98.0%. HTML, React, and Vue ISR values are 38.9, 34.6, and 32.7, below Claude Opus 4.5 at 52.2, 39.7, and 46.9.
- React CSR is 71.0 versus 70.7 for Claude Opus 4.5, and Vue CSR is 77.1 versus 74.3. Strong individual-check performance therefore coexists with weaker complete-project success.
The table separates build success, individual-check completion, and complete-instance completion. GLM-5 reaches 100 BSR in React, Vue, and Svelte, but its frontend ISR remains below Claude Opus 4.5; repository exploration is slightly higher while chained-task performance is lower.
The judge first verifies that a project builds and launches, then reads code and tests the live UI through screenshots and interaction. The illustrated mismatch between the wheel outcome and prize name shows why build success alone cannot establish functional correctness.
6.2.2 Backend Evaluation uses 85 tasks across Python, Go, C++, Rust, Java, and TypeScript, including feature implementation, bug fixes, regression repair, and performance optimization. Each Dockerized task has 5–10 human-crafted unit tests, and Pass@1 requires all associated tests to pass.
- GLM-5 scores 25.8, compared with 19.6 for GLM-4.7 and 26.9 for Claude Opus 4.5.
- The similar scores for GLM-5 and Claude Opus 4.5 apply to this curated internal task set and its strict unit-test criterion.
6.2.3 Long-horizon Evaluation separates repository exploration from sequential development. GLM-5 achieves 65.6 Pass@1 on repository exploration, compared with 47.8 for GLM-4.7 and 64.5 for Claude Opus 4.5, but reaches 52.3 on chained tasks versus 61.6 for Claude Opus 4.5.
- Exploration questions avoid filenames and code identifiers, require one or two semantic reasoning hops, and target obscure files at least three directory levels deep. Success means reading the target file, averaged over three runs, not completing a code change.
- Task chains come from merged PRs with tests and 3–15 linear-history commits. Semantic grouping and patch triage separate agent-produced code, verification tests, and automatically applied configuration or fixtures.
- Each task’s changes persist into subsequent tasks, and cumulative test patches check earlier behavior for regressions. Reported chained-task Pass@1 is per individual task, rather than an all-tasks-success rate for complete chains.
- The authors attribute the remaining gap partly to mistakes that compound through later development stages; the table does not isolate this cause experimentally.
6.2.4 Evaluation on evolving SWE tasks uses the January 2026 SWE-rebench set to supplement static SWE-bench Verified. GLM-5 resolves 42.1% with a reported SEM of ±1.21% and Pass@5 of 50.0%, while GLM-4.7 resolves 41.3% with SEM ±2.12% and Pass@5 of 56.3%.
- Claude Opus 4.5 resolves 43.8%, GPT-5.2 (xhigh) 51.7%, and Claude Opus 4.6 52.9%.
- The fresh-task evaluation extends the evidence beyond a longstanding public test set, but the small resolved-rate difference and lower Pass@5 do not establish a broad generalization advantage over GLM-4.7.
GLM-5’s January 2026 resolved rate of 42.1% is close to GLM-4.7 at 41.3% and Claude Opus 4.5 at 43.8%. Its Pass@5 of 50.0% is below GLM-4.7 at 56.3%, qualifying an across-the-board improvement claim on fresh SWE tasks.
6.3 Evaluation of Real-world General Abilities
Real-world general-ability evaluation combines internal human and automated tests with external assessments of translation, multilingual dialogue, instruction following, factual knowledge, and tool calling. Figure 11 reports improvements over GLM-4.7 on every displayed measure, including a rise from 60.8 to 95.8 on ToolCall-Badcase.
- Machine Translation uses 1,220 ZMultiTransBench samples across Zh to Es (300), Ru (250), Fr (220), Ko (200), Ja (150), Ar (50), and De (50), plus 753 English–Chinese MENT pairs. GPT-4.1 evaluates pairwise quality against a fixed baseline; displayed ZMultiTransBench and MENT-SNS scores rise from 1016 to 1050 and 993 to 1013. The figure also shows translation Human Scoring increasing from 2.85 to 3.12, without a corresponding detailed protocol in this subsection.
- Multi-lingual Dialogue combines LMArena Elo, increasing from 1441 to 1452, with 141 ZMultiDialBench instances scored by humans on a 1–10 scale, increasing from 8.44 to 8.57.
- Instruction Following uses 450 curated IF-Badcase instances, IF-Bench, and MultiChallenge. Displayed scores increase from 78.5 to 83.2, 68 to 72, and 61.9 to 64.8, respectively.
- World Knowledge scores increase from 31.0 to 36.9 on SimpleQA and from 72.9 to 75.2 on Chinese SimpleQA. Tool Calling uses 200 reviewed ToolCall-Badcase cases with verifiable tool and argument targets.
Every displayed general-ability measure improves relative to GLM-4.7, including translation Human Scoring from 2.85 to 3.12 and ToolCall-Badcase from 60.8 to 95.8. Different scales, dual axes, and evaluation protocols mean bar heights and apparent effect sizes cannot be compared directly across panels.
7 Conclusion
The conclusion presents GLM-5 as an open-weight model combining reasoning, efficient attention, and autonomous engineering workflows. The evaluations show progress over earlier GLM releases, while the internal complete-task metrics make clear that parity with proprietary systems depends on the workload and success criterion.
8 Easter Eggs
The Easter Eggs section identifies the anonymous OpenRouter release “Pony Alpha” as GLM-5 and recounts community reactions to its coding, agentic, and roleplay behavior. Preliminary attribution statistics report 25% guessing Claude Sonnet 5, 20% DeepSeek, 10% Grok, and the remainder GLM-5. Without a sample size or controlled comparison, these reactions are anecdotal release evidence rather than a capability evaluation.
9 Contribution
The Contribution section separates Core Contributors, Contributors, Tech Leads, and Advisors. It states that contributors’ names are listed in alphabetical order by first name.
Core Contributors
Core Contributors lists the principal contributor group, beginning with Chendi Ge and ending with Zixuan Li. The report provides names rather than individual technical-role assignments.
Contributors
Contributors records the broader contributor group, beginning with Bojie Wang and ending with Zuzhou Zhang. No per-person mapping to architecture, data, training, or evaluation work is specified.
Tech Leads
Tech Leads names Aohan Zeng, Xin Lv, Zhenyu Hou, Zhengxiao Du, Qinkai Zheng, Bin Chen, and Da Yin.
Advisors
Advisors names Jie Tang, Yuxiao Dong, Juanzi Li, Hongning Wang, Minlie Huang, and Bin Xu.
Appendix
- A Hyper-Parameters specifies Muon, cosine decay, and batch-size warmup. The pre-training learning rate warms from 0 to 2e-4 and decays to 4e-5; mid-training decreases it linearly to 1e-5. DSA warmup decreases from 5e-3 to 2e-4, and sparse adaptation uses a constant 1e-5. Table 10 reports 3 dense layers, 75 MoE layers, 1 MTP layer, hidden dimension 6144, 64 attention heads, 256 experts, 8 routed experts, 1 shared expert, and vocabulary size 154880; parameter counts include MTP but exclude embeddings and the output layer. The listed layer counts do not directly reconcile with the main text’s reported 80 layers.
- B Evaluation Details and B.1 Evaluation of Base Models qualify the final-model gains. GLM-5-Base improves EvalPlus Pass@1 to 87.0 from 78.1 for GLM-4.5-Base and LiveCodeBench-Base to 34.4 from 28.1, but GSM8K falls from 79.4 to 68.8 and MATH from 61.0 to 56.4. English and Chinese results are also mixed: MMLU rises from 86.1 to 88.3 and C-Eval from 86.9 to 88.8, while PIQA falls from 85.3 to 84.6 and C3 from 83.1 to 80.3.
- B.2 Evaluation of ARC Benchmarks supplies decoding and resource conditions: reasoning uses temperature = 1.0, top_p = 0.95, max_new_tokens = 131072; SWE uses temperature = 0.7, top_p = 0.95, max_new_tokens = 16384 and a 200K context window. Terminus-2 uses timeout = 2h, temperature = 0.7, top_p = 1.0, max_new_tokens = 8192, a 128K context, 16 CPUs, and 32 GB RAM. Claude Code 2.1.14 Terminal-Bench evaluation uses temperature = 1.0, top_p = 0.95, max_new_tokens = 65536, removes wall-clock limits, retains per-task resource limits, and averages five runs. CyberGym uses Claude Code 2.1.18 without web tools, temperature = 1.0, top_p = 1.0, max_new_tokens = 32000, a 250-minute timeout, and single-run Pass@1 over 1,507 tasks. BrowseComp retains the most recent 5 turns without full resets, while its context-management condition is described as discard-all; Vending-Bench 2 runs are conducted independently by Andon Labs.
- B.3 Optimized User Simulator for τ²-Bench documents Telecom and Retail prompt changes intended to prevent premature termination and preserve stated user constraints. B.4 Evaluation of Real-world Agentic Engineering Experience and B.4.1 Frontend Evaluation describe 220 frontend tasks across seven scenarios, with 113 HTML, 58 React, and 49 Vue tasks and 490, 249, and 210 check-items, respectively. Senior-expert task design, expert-reviewed checklist synthesis, execution-based correction, and removal of trivial tasks underpin benchmark construction; automated scoring therefore still depends on substantial human curation and validation.
Brief Thoughts
The report connects model training to concrete requirements of long-running agents: exact token alignment, consistent sparse selection, cache-local routing, failure-aware sampling, and executable feedback. Its engineering evaluation also separates successful builds and individual checks from complete tasks: high frontend BSR does not eliminate ISR gaps, and stronger repository exploration does not eliminate chained-development failures. Sparse-attention adaptation has direct quantitative evidence from matched-data SFT curves and long-context tables, although the efficient-attention comparisons do not share a single backbone and budget. Asynchronous-efficiency and hardware-cost claims are harder to assess because matched throughput, utilization, and deployment-cost measurements are absent. Comparisons require the stated conditions: text-only versus full HLE, original versus verified Terminal-Bench, modified simulators and timeouts, and the unresolved BrowseComp context-management description. Fresh SWE-rebench results and base-model math regressions limit an across-the-board improvement claim. The measured gains do not yet establish reliable end-to-end engineering.