发现量化后大模型性能可预测的规律,让压缩更可靠。
Scaling Laws for Post Training Quantized Large Language Models

- 通过实验分析量化后模型的局部损失曲面特征
- 提出统计模型能准确预测不同精度下的模型表现
- 适合做模型压缩与部署的工程师参考
已知大规模语言模型在预训练后的泛化能力随模型规模呈可预测的缩放关系。然而,与预训练阶段的实用缩放规律相比,后训练量化后的模型质量仍难以预测,实践中常需逐例验证。本文针对后训练权重量化,对多种大模型族在多种低精度张量数据类型下使用主流量化技术进行了系统性实证研究。我们识别出影响性能的关键缩放因子,即局部损失景观的特性,并基于此构建统计模型,可较准确地预测量化后模型的表现。
原文摘要 · Abstract (English)
Generalization abilities of well-trained large language models (LLMs) are known to scale predictably as a function of model size. In contrast to the existence of practical scaling laws governing pre-training, the quality of LLMs after post-training compression remains highly unpredictable, often requiring case-by-case validation in practice. In this work, we attempted to close this gap for post-training weight quantization of LLMs by conducting a systematic empirical study on multiple LLM families quantized to numerous low-precision tensor data types using popular weight quantization techniques. We identified key scaling factors pertaining to characteristics of the local loss landscape, based on which the performance of quantized LLMs can be reasonably well predicted by a statistical model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。