学习率调度与凸优化理论惊人吻合,可指导大模型训练调参。
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
- 用凸优化理论推导常数学习率加线性退火的性能界。
- 理论界无对数项,解释了退火策略的实际收益。
- 基于理论优化学习率,提升124M/210M模型训练效果。
我们发现大模型训练中的学习率调度行为与非光滑凸优化理论中的性能界出人意料地一致。针对常数学习率加线性退火的调度,我们给出了一个理论性能界;其中,退火带来的实际优势体现在界中无对数项。进一步表明,这种理论与实践的惊人吻合可用于学习率调优:通过延长训练时的最优学习率调度,以及跨调度迁移最优学习率,我们在124M和210M的Llama型模型上实现了显著的训练改进。
原文摘要 · Abstract (English)
We show that learning-rate schedules for large model training behave surprisingly similar to a performance bound from non-smooth convex optimization theory. We provide a bound for the constant schedule with linear cooldown; in particular, the practical benefit of cooldown is reflected in the bound due to the absence of logarithmic terms. Further, we show that this surprisingly close match between optimization theory and practice can be exploited for learning-rate tuning: we achieve noticeable improvements for training 124M and 210M Llama-type models by (i) extending the schedule for continued training with optimal learning-rate, and (ii) transferring the optimal learning-rate across schedules.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。