arXiv:2508.01483cs.LGcs.AI2025-08被引 10

研究学习率调度的冷却阶段,发现其形状影响模型性能。

Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler

  • 分析冷却阶段不同下降形状对模型的影响。
  • 高β₂值在冷却期能持续提升性能,效果相当于调整形状。
  • 可视化损失景观支持河流山谷理论,指导超参优化。

学习率调度在Transformer训练中至关重要,最终的衰减阶段对取得最佳性能起关键作用。然而,冷却阶段特有的损失下降机制仍不明确。为此,本文聚焦于暖启动-稳定-衰减(WSD)学习率调度器中的冷却阶段,进行系统性分析。结果表明,不同冷却形状揭示了模型中的基本偏差-方差权衡,平衡探索与利用的形状表现最优。同时,调节AdamW超参数在冷却期也带来显著性能差异,与选择冷却形状相当。值得注意的是,冷却期使用更高的β₂值可获得一致改进。从损失景观视角,本文提供了冷却期的景观可视化,实证支持河流山谷损失观。这些发现为Transformer训练中配置WSD调度器提供实用建议,强调应像传统超参调优一样重视冷却阶段的优化。

原文摘要 · Abstract (English)

Learning rate scheduling is essential in transformer training, where the final annealing plays a crucial role in getting the best performance. However, the mechanisms behind this cooldown phase, with its characteristic drop in loss, remain poorly understood. To address this, we provide a comprehensive analysis focusing solely on the cooldown phase in the Warmup-Stable-Decay (WSD) learning rate scheduler. Our analysis reveals that different cooldown shapes reveal a fundamental bias-variance trade-off in the resulting models, with shapes that balance exploration and exploitation consistently outperforming alternatives. Similarly, we find substantial performance variations $\unicode{x2013}$ comparable to those from cooldown shape selection $\unicode{x2013}$ when tuning AdamW hyperparameters. Notably, we observe consistent improvements with higher values of $β_2$ during cooldown. From a loss landscape perspective, we provide visualizations of the landscape during cooldown, supporting the river valley loss perspective empirically. These findings offer practical recommendations for configuring the WSD scheduler in transformer training, emphasizing the importance of optimizing the cooldown phase alongside traditional hyperparameter tuning.

学习率调度Transformer超参优化损失景观

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。