arXiv:2411.04330cs.LGcs.CL2024-11ICLR被引 95

提出精度感知的缩放定律,预测低精度训练与推理的性能损失。

Scaling Laws for Precision

  • 用有效参数量解释低精度训练的性能下降机制。
  • 发现更多数据训练反而让低精度推理更差,可能适得其反。
  • 适用于大模型低精度部署,尤其适合追求计算效率的研究者。

低精度训练与推理影响语言模型的质量与成本,但现有缩放定律未考虑此因素。本文提出针对训练与推理的精度感知缩放定律。我们提出:低精度训练会降低模型的‘有效参数量’,从而可预测训练中低精度及后训练量化带来的额外损失。对于推理,发现随着预训练数据增多,后训练量化引入的退化加剧,最终使更多数据反而有害。训练方面,我们的缩放定律能预测不同精度组合下的损失,并建议在更低精度下训练更大模型可能更节省算力。我们统一了预训练与后训练量化的缩放规律,得到一个统一函数形式,可预测多精度下的训练与推理退化。模型在超过465次预训练运行上拟合,并在1.7B参数、最多260亿标记物的数据集上验证预测效果。

原文摘要 · Abstract (English)

Low precision training and inference affect both the quality and cost of language models, but current scaling laws do not account for this. In this work, we devise "precision-aware" scaling laws for both training and inference. We propose that training in lower precision reduces the model's "effective parameter count," allowing us to predict the additional loss incurred from training in low precision and post-train quantization. For inference, we find that the degradation introduced by post-training quantization increases as models are trained on more data, eventually making additional pretraining data actively harmful. For training, our scaling laws allow us to predict the loss of a model with different parts in different precisions, and suggest that training larger models in lower precision may be compute optimal. We unify the scaling laws for post and pretraining quantization to arrive at a single functional form that predicts degradation from training and inference in varied precisions. We fit on over 465 pretraining runs and validate our predictions on model sizes up to 1.7B parameters trained on up to 26B tokens.

低精度缩放定律训练优化量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。