arXiv:2510.06213cs.LG2025-10被引 9

训练参数影响模型量化效果,调参可提升大模型压缩质量

Training Dynamics Impact Post-Training Quantization Robustness

  • 分析320亿参数、15万亿词的训练轨迹,发现学习率衰减后量化误差上升
  • 验证集损失与量化误差在学习率下降后分离,不受数据量影响
  • 通过可控实验证明调参可改善百亿级模型的量化鲁棒性

尽管后训练量化被广泛用于高效部署大语言模型,但其鲁棒性的内在机制仍不清晰。我们对开源语言模型的训练轨迹进行了全面分析,涵盖最高达320亿参数和15万亿训练标记的数据,以准确评估训练动态与量化性能之间的关系。关键发现表明,大规模训练中的量化误差由学习率及其他训练超参数的复杂相互作用驱动。具体而言,一旦学习率开始衰减,验证损失与量化误差便出现分离,且基本不受训练数据规模的影响。为探究对训练动态的干预措施并识别能有利调节量化鲁棒性的特定配置,我们在受控实验中训练了高达1000亿标记的模型。结果挑战了‘增大数据集规模必然削弱量化效果’的假设,反而表明通过战略性地调整训练超参数,可在大规模场景下提升量化质量。

原文摘要 · Abstract (English)

While post-training quantization is widely adopted for efficient deployment of large language models, the mechanisms underlying quantization robustness remain unclear. We conduct a comprehensive analysis of quantization degradation across open-source language model training trajectories up to 32B parameters and 15T training tokens to accurately assess the relationship between training dynamics and quantization performance. Our key finding is that quantization errors in large-scale training runs are driven by a complex interplay between learning rate and other training hyperparameters. Specifically, once learning rates decay, validation loss and quantization error diverge, largely independent of training data scale. To investigate interventions on the training dynamics and identify specific configurations that can modulate quantization robustness favorably, we train our own models in controlled experiments up to 100B tokens. Our results challenge the assumption that increasing dataset scale inherently compromises quantization effectiveness, demonstrating instead that strategic training hyperparameter interventions can improve quantization quality at scale.

量化大模型训练动态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。