source: Hugging Face Blog: LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation

level: technical

Liquid AI released QAD Q4_0 GGUFs, updated 4-bit checkpoints for LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B. These checkpoints use quantization-aware distillation, where a high-precision teacher model is distilled into a quantized student model. The result is the same memory footprint and throughput as native Q4_0 GGUFs but with much less quality loss. Developers can now run LFM2.5 models on edge devices without the usual accuracy drop from 4-bit quantization.

Benchmarks across reasoning, instruction-following, tool use, and agentic tasks show QAD checkpoints retain 97.1%, 96.5%, 97.4%, and 96.6% of their BF16 baseline performance for the four model sizes. On real hardware, the 230M and 350M QAD Q4_0 checkpoints match Q5_K_M quality at 4-33% higher decode throughput, while the 1.2B and 2.6B checkpoints match Q4_K_M quality at 3-14% higher throughput. They also match Unsloth's UD-Q4_K_XL where applicable.

The checkpoints are available on Hugging Face and work with llama.cpp or any GGUF Q4_0 runtime. This release builds on Liquid AI's earlier edge-focused models, including LFM2.5-VL-3B for vision and LFM2.5-2.6B for local agents. Quantization-aware distillation offers a practical path to deploy capable language models on constrained devices like Raspberry Pi or smartphones, reducing the trade-off between model size and performance.

why it matters: Developers can deploy LFM2.5 models on edge hardware with near-BF16 accuracy at 4-bit memory and speed, enabling more capable on-device AI.


source: Hugging Face Blog: LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation