Z.ai launches GLM-5.3-Flash with hybrid sparse-linear attention and 18B active parameters
Z.ai released GLM-5.3-Flash, a 320B-parameter multimodal model with only 18B active parameters using hybrid sparse-linear attention. The MIT-licensed model approaches Claude Opus 4.8 on coding benchmarks at one-tenth the inference cost of its predecessor.