提出损失动态的函数缩放律,揭示学习率调度对训练效率的影响。
Functional Scaling Laws in Kernel Regression: Loss Dynamics and Learning Rate Schedules
- 用内在时间视角分析核回归中的梯度下降,捕捉完整损失轨迹。
- 在数据和算力受限下,证明模型容量越大越高效,学习率衰减提升训练效率。
- 适用于大模型训练预测,解释温升-稳定-衰减调度优于纯衰减现象。
缩放律已成为理解与指导大语言模型(LLM)训练的统一视角。然而现有研究主要关注最终步损失,未明确整个损失动态是否遵循类似规律,以及学习率调度(LRS)如何塑造这些规律。本文在受控理论框架下,分析功率律核回归中随机梯度下降(SGD)的训练过程。关键洞察是引入新颖的内在时间视角,比迭代次数更准确刻画训练进度。我们建立函数缩放律(FSL),可描述任意学习率调度下的完整损失轨迹,其调度影响通过简单卷积函数体现。进一步针对三种典型调度——常数、指数衰减、温升-稳定-衰减(WSD)——推导出数据与算力受限场景下的显式缩放关系。这些结果解释了关键经验现象:(i) 容量更高的模型具有更高的数据与算力效率;(ii) 学习率衰减能提升训练效率;(iii) WSD型调度优于纯衰减。最后,在0.1B至1B参数的LLMs上实验表明,FSL作为大规模预训练中损失轨迹的代理模型具有实际预测能力。
原文摘要 · Abstract (English)
Scaling laws have emerged as a unifying lens for understanding and guiding the training of large language models (LLMs). However, existing studies predominantly focus on the final-step loss, leaving open whether the entire loss dynamics obey similar laws and, crucially, how the learning rate schedule (LRS) shapes them. We address these gaps in a controlled theoretical setting by analyzing stochastic gradient descent (SGD) on a power-law kernel regression model. The key insight is a novel intrinsic-time viewpoint, which captures the training progress more faithfully than iteration count. We then establish a Functional Scaling Law (FSL) that captures the full loss trajectory under arbitrary LRSs, with the schedule's influence entering through a simple convolutional functional. We further instantiate the theory for three representative LRSs -- constant, exponential decay, and warmup-stable-decay (WSD) -- and derive explicit scaling relations in both data- and compute-limited regimes. These comparisons explain key empirical phenomena: (i) higher-capacity models are more data- and compute-efficient; (ii) learning-rate decay improves training efficiency; and (iii) WSD-type schedules outperform pure decay. Finally, experiments on LLMs ranging from 0.1B to 1B parameters demonstrate the practical relevance of FSL as a surrogate model for fitting and predicting loss trajectories in large-scale pre-training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。