4比特量化模型在可靠性上表现最佳,兼顾效率与稳定性。
Reliability Scaling Laws for Quantized Large Language Models

- 研究4至8比特量化下模型的不确定性、校准性与鲁棒性。
- 4比特量化时可靠性达峰值,优于更高或更低比特精度。
- 量化反而提升模型对真实输入扰动的鲁棒性,适合部署场景。
量化通过降低参数位宽,使大语言模型在资源受限条件下仍具强大能力。尽管量化模型在未扰动输入上表现优异,其在扰动输入下的可靠性仍缺乏系统评估。本文全面评估了2、3、4、8比特下六种量化方法的可靠性,包含:(1) 不确定性:使用标准度量评估不同量化方案的可信度;(2) 校准性:分析模型规模与位精度对不确定性估计校准程度的影响;(3) 鲁棒性:设计字符级与词级输入扰动,测试模型在语义保持变化下的可靠性。结果表明,性能随总位数单调上升,但可靠性呈非线性变化,4比特量化模型达到可靠性峰值,展现出最优的可靠性-效率权衡。此外,实证发现量化能增强模型对自然输入扰动的鲁棒性。
原文摘要 · Abstract (English)
Quantization is a powerful strategy to build capable and resource-efficient large language models (LLMs) by reducing the bitwidth of the parameters. While quantized LLMs achieve state-of-the-art performance on unperturbed inputs using standard predictive metrics, their performance on perturbed inputs, measured using reliability metrics, remains underexplored, despite its importance for reliable deployment. To address this gap, we first conduct a comprehensive reliability evaluation of quantized LLMs consisting of three key components: (1) Uncertainty: We assess the trustworthiness of LLMs quantized to 2, 3, 4, and 8 bits using six different quantization methods, employing established uncertainty metrics. (2) Calibration: We assess how well-calibrated the uncertainty estimates of quantized models are across model scales and bit precisions. (3) Robustness: We design character-level and word-level input perturbations to evaluate the reliability of quantized models under semantically-preserving variations in the inputs that arise in real-world applications. Second, we characterize how reliability scales with the total number of model bits. Our study reveals that while the performance scales monotonically with the total number of bits, the reliability scalings are nonlinear. A reliability peak occurs for 4-bit quantized models, indicating that quantizing moderately sized models offers the best reliability-efficiency trade-off. Additionally, our empirical findings reveal that quantization enhances the robustness of LLMs to natural input perturbations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。