Z.ai released GLM-5.3 on 14 August 2026 with a blunt engineering claim: it is still the GLM-5.2 base. The company says every reported gain came from scaling post-training — more long-horizon environments, more RL compute on the slime stack — not a new pretrain. The pitch is twofold: stronger coding agents, and what Z.ai calls emergent cyber capability for vulnerability discovery and exploitation chains.
That matters because “.3” drops are easy to dismiss as marketing. Here the vendor is arguing the opposite: keep the foundation fixed, spend the month on post-train scale, and show the delta on public coding and cyber benches plus a real disclosure ledger.
Confirmed (vendor primary + docs)
- Primary launch post: GLM-5.3: Frontier Coding with Emergent Cyber Capabilities (Z.ai, 14 Aug 2026).
- Same base model as GLM-5.2; gains attributed entirely to post-training / environment-scaled RL.
- Vendor coding deltas vs GLM-5.2: ~50% on in-house Z.ai Code Bench; Terminal Bench 3.0 28.3 vs 4.6; DeepSWE v1.1 66.9 vs 46.2; Agents’ Last Exam 28.5 vs 23.8. Z.ai also claims open-source SOTA on Terminal Bench 3.0 and Agents’ Last Exam.
- On Z.ai Code Bench charts in the launch post, GLM-5.3 is shown beating Claude Opus 4.8 on some high-effort / token-efficiency points while remaining behind Claude Fable 5 at Max effort — still a vendor figure, not an independent audit.
- Cyber rows (vendor table): CyberGym 84.5% (vs 77.2% on 5.2; ahead of Mythos 5 / GPT-5.6 Sol on that row); ExploitBench 54.4% vs 24.4%; ExploitGym 105 / 130 tasks at 2h / 6h vs 29 / 39. Footnotes say ExploitGym budgets are throughput-normalized. Closed frontier models still lead by a wide margin on deeper exploit benches.
- Disclosure program: after review with Chinese security partners, 2,436 findings across 269 projects (1,097 medium/high); 53 public, 2,383 embargoed; tracked in the Z.ai Security Disclosure Ledger.
- Live now for GLM Coding Plan (Max / Pro / Lite) and ZCode; model IDs glm-5.3 and glm-5.3[1m]. Docs say older Coding Plan routing moves to the new engine.
- Open weights promised in ~two weeks after safety evaluation and hardening — not day-one Hugging Face weights.
- API/effort note (Coding Plan docs): default effort is max; levels are low / high / max. Requests that disable thinking (thinking.type: disabled and equivalents) are mapped to low and still run with lightweight thinking — they do not switch models. Explicit effort overrides the thinking toggle.
What's new
A shipped coding/cyber post-train on a frozen 5.2 base, live in Coding Plan today, with open weights staged behind safety review — and a public vuln ledger attached to the launch story, not only a leaderboard slide.
Analysis
The strongest claim is methodological, not the biggest number on the table: fixed base + scaled post-training. If outsiders can reproduce even part of the Terminal Bench 3.0 and Code Bench jumps, that strengthens the 2026 open-weight playbook where Chinese labs compete by iteration speed and RL environments rather than only by pretrain size.
The cyber narrative is sharper — and riskier. Z.ai says capability “emerged” faster than expected as vuln data entered the mix, with the largest gains further up the exploitation chain — exactly where it also remains furthest behind Mythos 5 and GPT-5.6 Sol. That is a candid gap chart, not a clean crown. CyberGym SOTA is interesting; ExploitGym/ExploitBench still read as “closing fast from behind.”
Product reality for builders is simpler than the scoreboard: Coding Plan already serves 5.3; weights are not free to self-host yet; migration defaults to max effort (disabled thinking maps to low, not off). Quota moved toward a points system with peak-hour (UTC+8 weekday afternoons) pricing — so “cheaper Chinese model” still needs a bill check, not a slogan.
Unknown
- Independent reproduction of Terminal Bench 3.0, Code Bench, and cyber harness numbers under Z.ai’s published footnotes.
- Exact open-weight license, card, and ship date when the two-week clock ends.
- How much of the disclosure haul transfers outside partner workflows and embargo rules.
- Public pay-per-token API list price for glm-5.3 versus Coding Plan points economics.
- Whether a later 5.5-class foundation drop still arrives, or near-term attention stays on post-train .x releases.
Our take
Believe the product event first: GLM-5.3 is live, same base as 5.2, weights delayed for safety. Treat the leaderboard as a vendor brief until third parties re-run it — especially cyber, where Z.ai both claims a CyberGym lead and shows a wide closed-frontier gap on deeper exploit. The story worth watching is not another version number; it is whether post-train-only jumps this large hold up outside Z.ai’s own harnesses.
Series: 1. GLM-5.2: What Actually Shipped and Why It Matters · 2. The 2026 Roadmap: Zhipu's Release Cadence and Strategy · 3. GLM-5.5 Rumors vs. Official Silence: Parsing the August 2026 Reports · 4. Developer's Guide: What to Treat as Confirmed vs. Unknown in the GLM Ecosystem · 5. GLM-5.3 Ships: Post-Training Coding and Emergent Cyber Defense · Zhipu GLM