source: Hugging Face Blog: Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original

level: research

researchers introduced quantization-aware healing, a method that recovers accuracy in models that have been both structurally compressed and quantized to 4 bits. the approach distills directly from the original, full-size, full-precision model instead of from the recovered checkpoint. this removes the accuracy ceiling imposed by the recovered model. applied to a gpt-oss 120b model compressed to 60b and quantized to mxfp4, the resulting model beat its own bfloat16 version on 7 of 9 benchmarks.

the 4-bit model gained 7.4 points on long-context reasoning and 5.6 points on aime 2025 math over the bfloat16 checkpoint. it also surpassed the full-size 120b teacher on livecodebench and came within 1.6 points on gpqa diamond. compared to quantization-aware training, the new method reached peak accuracy about 7 times faster and stayed stable, while qat lost nearly 19 points after its peak. the two methods tied on best-case accuracy at roughly 54.9 versus 54.6.

the method uses a chunked kl-divergence loss to handle long contexts up to 32k tokens without exceeding gpu memory. because the teacher is frozen, the student has no incentive to drift once it matches the teacher, unlike cross-entropy training. this makes the healing process more stable and less dependent on early stopping. the result suggests quantization can be an opportunity to improve a model rather than just a cost of efficiency, especially for deployment on smaller hardware.

why it matters: this method lets teams deploy smaller, cheaper 4-bit models that can outperform their full-precision versions, reducing serving costs without sacrificing accuracy.


source: Hugging Face Blog: Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original