arXiv:2608.24814cs.LGstat.ML2026-08被引 1

学习率与参数范数的比值决定语言模型训练损失变化,是统一控制变量。

Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining

  • 用有效学习率(ELR)统一不同训练设置下的损失轨迹。
  • 跨优化器、架构、数据集,损失轨迹匹配误差仅约0.001~0.003。
  • 适合研究训练稳定性、超参调优或模型缩放规律的研究者。

我们发现语言模型预训练中存在有效学习率(ELR)坍缩现象:学习率(LR)与参数范数主要通过其比值——有效学习率(ELR)——共同决定损失动态。当不同实验的ELR相同时,尽管学习率和参数范数差异显著,其损失轨迹仍高度重合,平均坍缩误差通常为几×10⁻³,低于典型配置下的种子间波动。系统性消融分析表明,归一化设计及LR-范数变化的时间尺度是影响坍缩精度的关键因素。受控干预显示,权重衰减与超球形损失函数主要通过其诱导的ELR调度影响训练动态。以ELR替代学习率后,可实现拟合函数缩放律(FSL)在不同范数控制方法间的迁移。由此构建的基于ELR的FSL还能解释范数控制带来的延迟加速现象。这些结果确立了ELR作为连接学习率调度、范数控制与损失动态的通用坐标。

原文摘要 · Abstract (English)

We uncover ELR collapse in language model pretraining: learning rate (LR) and parameter norm govern loss dynamics primarily through their ratio, the effective learning rate (ELR). When ELR is matched across runs, their loss trajectories collapse throughout training despite substantially different LRs and parameter norms. Across optimizers, architectures, datasets, and model scales, mean collapse errors are typically a few x 10^-3, below the seed-to-seed variation measured in a representative configuration. Systematic ablations identify normalization design and the timescale of LR-norm variation as key determinants of collapse precision. Controlled interventions further show that weight decay and Hyperball shape loss dynamics primarily through the ELR schedules they induce. Replacing LR with ELR enables a fitted functional scaling law (FSL) to transfer across norm-control methods. The resulting ELR-based FSL also explains delayed acceleration, a recurring effect of norm control. Together, these results establish ELR as a common coordinate linking LR scheduling, norm control, and loss dynamics.

训练动态有效学习率语言模型缩放律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。