Skip to content
Main Site News Console

Quantization-Aware Repair: A Compressed 4-Bit Model That Outperforms the Full-Precision Original

· Hugging Face Translated
教程模型卡

Making a large language model smaller almost always comes at a cost. Today, efficient deployment typically follows this approach: first compress the architecture by removing layers, attention heads, or neurons to reduce the parameter count; then further quantize the remaining weights to 4 bits to reduce memory and computational overhead. Both steps can deliver significant savings, but when combined, they systematically weaken the capabilities people actually care about: reasoning, mathematical problem-solving, and code generation. As a result, serious deployment workflows typically add a recovery step before putting a model into production, commonly referred to as “healing.” Recently released open-weight models, including gpt-oss, NVIDIA’s Nemotron series, and our own Hypernova 60B, all use some form of “compress first, heal later” approach.

Our latest paper, Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs, investigates a long-standing open question in the field: how much does this recovery step actually help when a model has undergone not only quantization but also structural compression, and what is the right way to implement it? We propose Quantization-Aware Healing (QAH). Applied to a compressed GPT-OSS 120B model whose parameter count was reduced from 120B to 60B and quantized to MXFP4, QAH produced a model that outperformed its own full-precision (bfloat16) version on 7 of 9 benchmarks. This 4-bit model is smaller, cheaper to run, and more accurate than the original checkpoint used for quantization. This overturns the usual relationship between a 4-bit model and the 16-bit model from which it originated.

Why Conventional Healing Falls Short Here

Most efficiency optimization pipelines follow the same three steps: compress the architecture, quantize the compressed weights, and then heal the resulting damage. The difference between approaches lies entirely in the final step.

The dominant healing approach today is quantization-aware training (QAT). It inserts fake-quantization operators into the forward pass and continues fine-tuning the model with task loss, allowing the weights to learn to adapt to the lower-precision representation. In practice, this means re-running the already expensive, multi-stage post-training process—including supervised fine-tuning, RLHF, and agent tuning—through noisier, lower-precision forward passes. This is costly; as our results show, training can also become unstable if it continues for too long after reaching its optimum.

Another approach is quantization-aware distillation (QAD), which avoids repeating the earlier training process. Rather than using task loss, QAD directly distills a frozen, full-precision teacher model into a quantized student model by computing a KL divergence loss over the output logits. This works well when quantization is the only change, because there is a genuinely identical full-precision version of the model available to serve as the teacher.

But this assumption breaks down when the model undergoes structural compression—removing layers, attention heads, or neurons rather than merely reducing the number of bits. There is no independently trained full-precision version of the smaller architecture. The only available teacher candidate is the healed bfloat16 checkpoint, which is itself a distilled approximation of the original model. Distilling from it constrains the quantized student to an already degraded target and caps its accuracy at the level of the healed checkpoint itself.

Thus, how to heal a model that has undergone both structural compression and quantization has remained a genuinely open problem until now.

Our Approach

QAH removes this ceiling through one change: instead of distilling from the healed model, it distills directly from the original model before compression. The teacher and student models do not even share the same architecture. The teacher is full-sized and operates at full precision; the student is half the size and runs in MXFP4. Because the teacher’s output distribution is architecture-independent, a mismatch in model size or shape does not prevent knowledge transfer. The student never sees hard labels; it only receives the teacher’s output distribution and matches it through KL divergence over the logits.

This redefines the role of quantization. In QAH, quantization is no longer a lossy post-processing step performed after healing. Instead, it becomes a second complete round of distillation from the original teacher—a source of supervision that the bfloat16 checkpoint never received. The 4-bit student is not compensating for information lost through quantization; it is acquiring information that the earlier healing stage failed to transfer because of limited time or data.

The loss function itself also provides stability benefits. Because KL distillation anchors the student to a fixed teacher distribution, once the student catches up with the teacher, it is not pressured to drift further. By contrast, cross-entropy task loss continues pushing the student toward hard labels and does not stop on its own. This difference matters both for accuracy and training stability, as the comparison below demonstrates.

To enable QAH to handle long-context scenarios—the healing corpus contains documents up to 32k tokens long—we reused the memory-efficient chunked KL-divergence loss from the efficient distillation approach introduced in the accompanying paper. This loss computes KL divergence one sequence slice at a time without materializing the complete “vocabulary × sequence” grid in memory, allowing 32k-token healing to fit within a fixed GPU memory budget. We described the mechanics of this loss in detail in our previous article.

QAH overview: Model capabilities decline substantially after structural compression and quantization. QAH distills from the original pre-compression model, using it as a frozen teacher, without requiring the multi-stage post-training process to be rerun in order to recover performance.

QAH overview. Model capabilities decline substantially after structural compression and quantization. QAH distills from the original model: the original model serves as the frozen teacher, whose logits are precomputed offline, rather than distilling from the healed checkpoint. Source: Figure 1 of the paper.

Results

We applied QAH to a GPT-OSS 120B model: first compressing it to 60B parameters and healing it in bfloat16, then requantizing it to MXFP4 during the QAH process. The most natural comparison is the bfloat16 checkpoint of the same 60B model—the best full-precision version currently available for this architecture. The QAH model achieved better results on 7 of 9 benchmarks.

Benchmark120B Teacher (MXFP4)60B BF16 (Healed)60B MXFP4 (QAH)QAH vs. BF16
AA-LCR (Long-Context Reasoning)50.035.342.7+7.4
AIME 2025 (Mathematics)80.070.776.3+5.6
Aider (Agentic Coding)45.338.240.9+2.7
τ²-bench (Tool Use)68.459.461.7+2.3
GPQA Diamond (Science)69.065.767.4+1.7
IFBench (Instruction Following)63.358.459.9+1.5
LiveCodeBench (Coding)66.065.566.5+1.0
MMLU-Pro (Knowledge)78.074.073.8−0.2
SciCode (Scientific Coding)37.535.634.2−1.4

The two benchmarks on which QAH fell behind were MMLU-Pro and SciCode, with gaps of less than 1 point and 1.5 points, respectively. On every other test, the 4-bit model outperformed its own 16-bit source model. The largest gains appeared precisely in the capabilities most heavily damaged by structural compression: long-context reasoning (a 7.4-point improvement on AA-LCR) and mathematical ability (a 5.6-point improvement on AIME 2025).

The comparison with the original 120B teacher model is equally instructive. Although the QAH model has half as many parameters as the teacher and roughly one-quarter of its weight memory, it surpassed the full-sized teacher on LiveCodeBench (66.5 vs. 66.0) and fell behind by only 1.6 points on GPQA Diamond (67.4 vs. 69.0). The largest remaining gap with the teacher appeared on AA-LCR. This is an extreme long-context benchmark, where the capacity lost through compression is most pronounced.