arXiv:2504.09629cs.LGstat.AP2025-04NeurIPS被引 36

提出量化误差传播机制,显著提升低比特压缩下大模型精度

Quantization Error Propagation: Revisiting Layer-Wise Post-Training Quantization

  • 显式建模层间量化误差传播并动态补偿
  • 在极低比特(如2/3比特)下性能超越现有方法
  • 轻量可调,适配不同模型与资源约束

层间后训练量化(Layer-wise PTQ)因无需微调而成为压缩大语言模型的有力手段,但近期进展趋于停滞,亟需重新审视其核心瓶颈。本文发现现有方法在低比特条件下,量化误差逐层累积导致性能严重下降。为此,提出量化误差传播(QEP)框架,通过显式传播并补偿累积误差,提升层间量化鲁棒性。QEP具备可调的传播机制,既能防止过拟合,又控制计算开销,支持多种模型架构与资源预算。在多个大语言模型上的实验表明,QEP增强的层间PTQ显著优于现有方法,尤其在极端低比特(如2/3比特)场景下优势明显。

原文摘要 · Abstract (English)

Layer-wise PTQ is a promising technique for compressing large language models (LLMs), due to its simplicity and effectiveness without requiring retraining. However, recent progress in this area is saturating, underscoring the need to revisit its core limitations and explore further improvements. We address this challenge by identifying a key limitation of existing layer-wise PTQ methods: the growth of quantization errors across layers significantly degrades performance, particularly in low-bit regimes. To address this fundamental issue, we propose Quantization Error Propagation (QEP), a general, lightweight, and scalable framework that enhances layer-wise PTQ by explicitly propagating quantization errors and compensating for accumulated errors. QEP also offers a tunable propagation mechanism that prevents overfitting and controls computational overhead, enabling the framework to adapt to various architectures and resource budgets. Extensive experiments on several LLMs demonstrate that QEP-enhanced layer-wise PTQ achieves substantially higher accuracy than existing methods. Notably, the gains are most pronounced in the extremely low-bit quantization regime.

量化大模型压缩误差补偿

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。