发现大模型权重量化最低可达1.58比特,否则表达能力会崩溃
On the Expressive Power of Weight Quantization in Large Language Models

- 从理论证明1.58比特是权重量化下限
- 比特数减少时模型表达力呈多项式下降
- 为模型压缩与加速提供理论依据,适合做量化研究者
近年来,将大语言模型的可学习参数以n比特格式进行权重量化,因其在模型压缩和推理加速方面的潜力而受到广泛关注。尽管已发展出诸多实用技术,但关于量化比特数减少时模型近似能力与表达力退化的理论理解仍不清晰。本文对大语言模型在不同量化比特数下的表达能力进行了理论研究,提出1.58比特是权重量化极限,并通过建立普遍逼近性与表达力坍塌特性加以证明。同时验证了权重量化会导致表达力退化,其容量随比特数减少呈多项式下降。这些理论发现为基于缩放定律的权重量化研究提供了坚实基础,并为未来模型压缩与推理加速研究提供了洞见。
原文摘要 · Abstract (English)
In recent years, weight quantization that encodes the learnable parameters of large language models in an $n$-bit format has garnered significant attention due to its potential for model compression and inference acceleration. Many practical techniques have been developed; however, the theoretical understanding of many aspects, especially the approximation and degradation of expressive power as the number of quantization bits decreases, remains unclear. In this paper, we provide a theoretical investigation into the expressive capability of large language models relative to the number of quantization bits. We argue that 1.58-bit is the limiting precision for weight quantization by establishing the universal approximation and expressive collapse properties of weight-quantized models with respect to the number of quantization bits. Additionally, we confirm that weight quantization leads to expressive degradation, in which the expressive capacity of weight-quantized models degrades polynomially as the number of quantization bits decreases. These theoretical findings provide a solid foundation for advancing weight quantization in the context of scaling laws and shed insights for future research in model compression and inference acceleration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。