arXiv:2411.17691cs.LGcs.CL2024-11被引 22

低比特量化下,训练不足的大模型表现更好,且可据此估算模型训练量。

Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens

  • 通过控制实验发现:模型越大或训练越少,低比特量化损伤越小。
  • 1500多个检查点实验证明,100万亿训练令牌下的大模型量化性能可能不理想。
  • 提出用量化损伤衡量模型训练水平,适合关注模型效率与部署的研究者。

我们发现,低比特量化对训练不足的大语言模型更有利:模型规模越大或训练令牌数越少,其在低比特量化下的性能退化(QiD)越小;而小模型即使经过大量训练,也面临显著的量化损伤。我们在受控环境下研究了超过1500个不同规模、不同训练程度(训练不足或完全训练)的量化模型检查点,推导出量化损伤与训练令牌数、模型大小和比特宽度之间的缩放定律。基于这些定律,我们提出一种新视角:可用量化损伤来评估模型的训练水平,并确定各类模型完全训练所需的训练令牌数。此外,利用缩放定律预测了在100万亿训练令牌下不同规模模型的量化表现,结果表明未来模型在低比特量化下的性能可能不理想。这为低比特量化带来潜在挑战,强调在评估量化研究时需关注模型的训练水平。为促进后续研究,我们已将所有1500+个量化检查点发布于https://huggingface.co/Xu-Ouyang。

原文摘要 · Abstract (English)

We reveal that low-bit quantization favors undertrained large language models (LLMs) by observing that models with larger sizes or fewer training tokens experience less quantization-induced degradation (QiD) when applying low-bit quantization, whereas smaller models with extensive training tokens suffer significant QiD. To gain deeper insights into this trend, we study over 1500 quantized LLM checkpoints of various sizes and at different training levels (undertrained or fully trained) in a controlled setting, deriving scaling laws for understanding the relationship between QiD and factors such as the number of training tokens, model size and bit width. With the derived scaling laws, we propose a novel perspective that we can use QiD to measure an LLM's training levels and determine the number of training tokens required for fully training LLMs of various sizes. Moreover, we use the scaling laws to predict the quantization performance of different-sized LLMs trained with 100 trillion tokens. Our projection shows that the low-bit quantization performance of future models, which are expected to be trained with over 100 trillion tokens, may NOT be desirable. This poses a potential challenge for low-bit quantization in the future and highlights the need for awareness of a model's training level when evaluating low-bit quantization research. To facilitate future research on this problem, we release all the 1500+ quantized checkpoints used in this work at https://huggingface.co/Xu-Ouyang.

量化大模型训练水平缩放定律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。