Tencent compresses 770B Hy4 MoE model to 214GB GGUF via layer-heterogeneous quantization
Tencent open-sourced an extreme quantized GGUF build of its 770B Hunyuan Hy4 model, compressing 1.5TB weights down to 214GB. By varying precision from 1.31-bit on non-critical layers to 2+ bits on sensitive layers, the model fits on accessible multi-GPU nodes.