从信号噪声比看大模型量化降损,揭示误差来源与传播规律
Quantization Degradation in Large Language Models: A Signal-Noise Perspective

- 用信号噪声比分析量化误差的产生与传播机制
- 3比特时性能下降明显,受任务类型和模型规模影响大
- 适合关注模型压缩与量化部署的研究者阅读
后训练量化可降低大语言模型的部署成本,但其性能下降程度并不仅由位宽决定。我们在多个模型家族上系统研究了不同位宽、量化方法、模型规模和下游任务下的纯权重后训练量化。发现4比特通常保持性能,2比特常导致广泛降损,而3比特时降损已显现,但差异显著取决于任务类型、量化方法和模型规模。为解释这一变异性,我们引入信噪比(SNR)来衡量量化扰动对全精度表示的影响。研究揭示降损源于两个关联过程:单个模块内量化误差的生成方式,以及跨层传递中的累积效应。首先,源信噪比分解表明新引入的误差依赖于权重误差幅度、任务特定信号强度,以及量化误差与任务激活的对齐程度,三者影响各异。其次,跨层传播分析显示,这些误差可在层间被衰减、保持或放大,且更大模型能更有效抑制误差放大。综合结果表明,量化降损由误差源头特性及网络中累积行为共同决定。
原文摘要 · Abstract (English)
Post-training quantization reduces the deployment cost of large language models, yet how severely a quantized model degrades is not determined by bit-width alone. We systematically study weight-only post-training quantization across bit-widths, quantization methods, model scales and downstream tasks on multiple model families. We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation, and at 3-bit, degradation becomes apparent but varies markedly with task type, quantization method and model scale. To explain this variability, we use the signal-to-noise ratio (SNR) to measure how strongly quantization perturbs full-precision representations. We trace degradation back to two linked processes: how quantization errors arise within individual modules, and how they accumulate across layers. First, a source SNR decomposition shows that newly introduced errors depend on three factors: the magnitude of the weight error, the strength of the task-specific signal, and how strongly the quantization error aligns with task-specific activations. Different factors affect these components in distinct ways. Second, a cross-layer propagation analysis shows that these errors can be attenuated, preserved, or amplified as they pass across layers, and that larger models benefit from weaker error amplification. Together, these results establish that quantization degradation is governed by how errors are introduced at the source and how they accumulate across the network.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。