Z.ai has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. Its hybrid architecture combines sparse and linear attention to sharply reduce long-context serving costs while preserving precise long-context capabilities. The model contains 320 billion total parameters with only 18 billion active during inference, delivering performance that outperforms GLM-5.2 across benchmarks at one-tenth the price.
The model adopts Manifold-Constrained Hyper-Connections (mHC) to improve scaling efficiency and was trained on a 30-trillion-token multimodal pre-training corpus. Z.ai claims GLM-5.3-Flash approaches Claude Opus 4.8 on coding and agentic benchmarks. Evaluation results published on the model card show performance on HLE with tools (full set), NL2Repo, DeepSWE, Terminal-Bench 2.1, Agent's Last Exam: Toolathlon Verified, AutomationBench, GDPval-AA v2, and BabyVision benchmarks.
What's new
- Hybrid sparse-linear attention architecture — first in the GLM series, designed to sharply reduce long-context serving costs
- Manifold-Constrained Hyper-Connections (mHC) — improves scaling efficiency across the model
- 320B total parameters, 18B active — sparse activation enables cost-efficient inference
- Native multimodal support — handles image-text-to-text tasks out of the box
- 30T-token multimodal pre-training corpus — latest training data for the GLM series
- MIT license — open weights available on Hugging Face
- 1M context window support — evaluated on NL2Repo with 1M context
- Multiple quantization formats — BF16, F8_E4M3, F32 available
Why it matters
The hybrid attention architecture represents a significant architectural shift for the GLM series, potentially setting a new efficiency frontier for large multimodal models. By activating only 18B parameters during inference while maintaining 320B total capacity, GLM-5.3-Flash could make frontier-class multimodal capabilities accessible at substantially lower compute costs. The MIT license and immediate availability on Hugging Face with support for Transformers, vLLM, SGLang, KTransformers, and Docker Model Runner lowers deployment barriers for researchers and developers. The model's evaluation on agentic benchmarks like DeepSWE and Terminal-Bench with 400K-600K context windows signals readiness for extended coding workflows.
Our take
GLM-5.3-Flash's sparse activation approach — 18B active from 320B total — is the most concrete efficiency claim in a field where "mixture-of-experts" labels often obscure actual active parameter counts. The hybrid attention design and mHC scaling improvements suggest Z.ai is targeting the long-context agentic workloads where current frontier models struggle with cost. Independent replication of the claimed Claude Opus 4.8 parity on coding benchmarks will be the real test.
Series: 1. Z.ai Ships GLM-5.2: Open-Weight Model Matches Opus 4.8 on Agentic Benchmarks at a Fraction of the Cost · 2. Z.ai launches GLM-5.3-Flash with hybrid sparse-linear attention and 18B active parameters · Zhipu GLM