Zhipu AI, now operating internationally as Z.ai, has compressed its model release cycle to roughly two-month intervals across the first half of 2026. The Beijing-based lab shipped GLM-4.5 in July 2025, followed by GLM-4.7-Flash in January, GLM-5 in February, GLM-5.1 in April, and GLM-5.2 in mid-June. Each iteration kept the 744-billion-parameter Mixture-of-Experts backbone with 40 billion active parameters while expanding context windows and adding agentic tooling.
The June 13 launch of GLM-5.2 arrived just two days after the U.S. Commerce Department ordered Anthropic to restrict foreign access to its Fable 5 and Mythos 5 models, a timing noted by multiple observers. Zhipu followed on July 2 with ZCode, a harness that turns GLM-5.2 into an autonomous coding agent with a 1-million-token context window and a novel goal-verification. The company also rolled out developer promotions including 50 percent higher data quotas for existing subscribers and 5 million free tokens for new ZCode users.
Confirmed
- GLM-4.5 released July 2025; first open model to run competitively on Cerebras hardware.
- GLM-4.7-Flash (30B MoE, 3B active) launched January 22, 2026; targets local coding on consumer hardware.
- GLM-5 (744B MoE, 40B active) launched February 12, 2026; trained on 28.5 trillion tokens using Huawei chips.
- GLM-5.1 released April 7, 2026; scored 58.4 on SWE-bench Pro, then the top open-source result.
- GLM-5.2 released June 13–16, 2026; 1-million-token context, MIT license, API at $1.40/$4.40 per 1M input/output tokens.
- ZCode agentic environment launched July 2, 2026; 173 tokens/second output, 1.4-second time-to-first-token.
- Zhipu listed on Hong Kong Stock Exchange January 2026, raising $558 million at ~$6.6B implied valuation.
- U.S. Entity List designation January 2025 accelerated domestic compute strategy and open-weight focus.
Analysis
The cadence reveals a deliberate strategy: hold the parameter budget constant while investing each cycle into context length, reasoning modes, and agentic infrastructure. GLM-5.1 to GLM-5.2 added a 5× context expansion (200K to 1M tokens), dual reasoning modes (High/Max), and IndexShare attention optimization that cuts per-token FLOPs by 2.9× at full context. The model architecture itself — 744B total, 40B active — has not changed since GLM-5 in February, suggesting Zhipu is amortizing the expensive pre-training run across multiple specialized checkpoints.
Pricing reinforces the volume play. The standalone API undercuts GPT-5.5 by roughly 6×, while the flat-rate GLM Coding Plan ($18–$160/month) targets high-throughput agentic workloads that burn output tokens. Peak-hour deductions at 3× and off-peak at 2× (temporarily 1× through September) align costs with the model's observed token appetite — independent testing by Artificial Analysis found GLM-5.2 consumes more tokens per task than rivals to reach comparable scores.
Benchmark positioning is nuanced. Vendor-reported SWE-bench Pro at 62.1 beats GPT-5.5's 58.6 but trails Claude Opus 4.8 at 69.2. Terminal-Bench 2.1 shows 81.0 versus Opus 4.8's 85.0 and GPT-5.5's 84.0. The one independent anchor — Artificial Analysis ranking GLM-5.2 #1 among open-weight models in its class (Intelligence Index 51 vs. median 25) — confirms leadership in the open tier while underscoring a persistent gap to the top closed models. Kimi K3's July 27 release (Intelligence Index 57) has since claimed the open-weight crown, illustrating how quickly the lead rotates.
Geopolitically, the Entity List designation and Anthropic's Fable/Mythos restriction created a window Zhipu exploited: an open-weight, MIT-licensed alternative available globally without U.S. export controls. The Hong Kong IPO provided capital to sustain the compute-intensive cadence while the Huawei-chip training run for GLM-5 demonstrated supply-chain independence.
Unknown
- No official 2026 roadmap beyond GLM-5.2 has been published by Zhipu; rumored GLM-5.5 remains unconfirmed.
- Whether the two-month cadence can continue given pre-training compute demands and the 28.5T token dataset ceiling.
- Long-term data-governance posture for international enterprise customers given China-based operations.
- Impact of Kimi K3's open-weight lead on Zhipu's developer adoption and pricing power.
Our take
Zhipu's rhythm looks less like a sprint and more like a sustained production line: freeze the expensive backbone, iterate the serving stack, and price for volume. The real test is whether the open-weight moat holds once Chinese rivals like Moonshot (Kimi K3) match the licensing and context advantages. For now, the 6× cost advantage on agentic coding is the strongest signal — if your pipeline burns tokens, the math works regardless of who sits atop the leaderboard next month.
Series: 1. GLM-5.2: What Actually Shipped and Why It Matters · 2. The 2026 Roadmap: Zhipu's Release Cadence and Strategy · 3. GLM-5.5 Rumors vs. Official Silence: Parsing the August 2026 Reports · 4. Developer's Guide: What to Treat as Confirmed vs. Unknown in the GLM Ecosystem · 5. GLM-5.3 Ships: Post-Training Coding and Emergent Cyber Defense · Zhipu GLM
Sources
- SCMP: Zhipu AI releases harness for GLM-5.2 model as Chinese firm takes aim at Anthropic
- OpsMatters: GLM-5.2 Review (2026): Zhipu AI's Open-Weight Coding Model, Honestly Assessed
- Fello AI: What Is GLM 5.2? Zhipu's 1M-Context Open Model
- ThursdAI: Zhipu AI (GLM) Releases: GLM-5.2, GLM-4.7-Flash & 9 More
- Turing Post: Z.ai, Formerly Zhipu AI: The Rise of an AI Tiger Reaching for AGI