arXiv:2505.07796cs.CLcs.AI2025-05ICML被引 7

揭示大模型持续预训练的学习规律,可预测任意阶段的损失表现。

Learning Dynamics in Continual Pre-Training for Large Language Models

  • 通过解耦分布偏移与学习率衰减,建立持续预训练的统一数学模型。
  • 提出缩放定律,能准确预测不同训练步数和学习率策略下的损失值。
  • 适用于定制化调参,平衡通用性能与下游任务表现,适配多场景应用。

持续预训练(CPT)已成为将强大基础模型应用于特定下游任务的流行且有效方法。本文研究了大语言模型在持续预训练过程中的学习动态,重点关注通用性能与下游领域性能随训练步骤的变化,以验证损失衡量领域表现。我们发现CPT损失曲线本质上表征了从一个曲线过渡到另一个隐藏曲线的过程,并可通过解耦分布偏移与学习率衰减的影响来描述。由此推导出结合两个因素的CPT缩放定律,能够预测任意(持续)训练步数及不同学习率调度策略下的损失。该公式全面揭示了CPT中的多个关键因素,包括损失潜力、峰值学习率、训练步数、重播比例等。此外,该方法可适应不同CPT目标,如平衡通用性与领域性能。大量实验表明,该缩放定律在多种CPT数据集和超参数设置下均成立。

原文摘要 · Abstract (English)

Continual Pre-Training (CPT) has become a popular and effective method to apply strong foundation models to specific downstream tasks. In this work, we explore the learning dynamics throughout the CPT process for large language models. We specifically focus on how general and downstream domain performance evolves at each training step, with domain performance measured via validation losses. We have observed that the CPT loss curve fundamentally characterizes the transition from one curve to another hidden curve, and could be described by decoupling the effects of distribution shift and learning rate annealing. We derive a CPT scaling law that combines the two factors, enabling the prediction of loss at any (continual) training steps and across learning rate schedules (LRS) in CPT. Our formulation presents a comprehensive understanding of several critical factors in CPT, including loss potential, peak learning rate, training steps, replay ratio, etc. Moreover, our approach can be adapted to customize training hyper-parameters to different CPT goals such as balancing general and domain-specific performance. Extensive experiments demonstrate that our scaling law holds across various CPT datasets and training hyper-parameters.

持续学习大模型训练缩放定律优化机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。