source: Hugging Face Blog: Transformers now runs llama.cpp quants

level: technical

hugging face added support for running gguf models in transformers. users can load a gguf checkpoint from the hub with from_pretrained and generate text on an apple silicon mac. the integration reuses llama.cpp's ggml metal kernels through the kernels library. initial support targets the qwen3.5 architecture. the goal is to make local inference practical with familiar transformers apis while keeping performance close to llama.cpp.

benchmarks on a macbook pro m2 max show transformers within a few percent of llama.cpp for three checkpoints. for unsloth qwen3.5-4b q4_k_m, transformers reached 41.2 tokens per second versus 42.1 for llama.cpp. a larger dense model and a mixture-of-experts model showed similar gaps. the transformers measurement includes prompt prefill, while llama-bench reports decode-only throughput. without compatible kernels, the loader falls back to dequantizing the model and uses more memory.

the work also improves generate for all transformers models by removing an unnecessary attention mask and deferring stopping checks. beyond gguf, the ggml kernels can accelerate models that llama.cpp does not support, including vision, audio, and multimodal architectures. users can serve gguf models through an openai-compatible api with transformers serve. the integration lets developers inspect activations, evaluate quantized checkpoints, and fine-tune from dequantized weights.

why it matters: data scientists can now run quantized models locally with transformers, making experimentation cheaper and more accessible on laptops.


source: Hugging Face Blog: Transformers now runs llama.cpp quants