Hugging Face adds GGUF support to Transformers for local inference on Apple Silicon
Hugging Face added GGUF support to Transformers, letting developers load llama.cpp checkpoints directly on Apple Silicon Macs. The integration reuses ggml Metal kernels and keeps the generation loop in PyTorch.