小模型可精准预测大模型的最佳学习率调度方案。
Scaling and Transferability of Annealing Strategies in Large Language Model Training
- 基于预热-平稳-衰减框架,建立可迁移的学习率优化模型。
- 不同规模模型的最优学习率衰减速率具有一致性模式。
- 无需反复调参,适合追求高效训练的AI研发人员。
学习率调度对大语言模型训练至关重要,但跨模型配置的最优退火策略仍不明确。本文研究了大模型训练中退火动态的可迁移性,改进了温升-平稳-衰减(WSD)调度器的通用预测框架。新框架融合训练步数、最大学习率和退火行为,实现更高效的调度优化。实验在密集模型与混合专家(MoE)模型上验证,表明最优退火比例在不同训练配置间具有稳定模式,可跨模型迁移。该方法为选择最优退火策略提供实用指导,避免繁琐超参数搜索。
原文摘要 · Abstract (English)
Learning rate scheduling is crucial for training large language models, yet understanding the optimal annealing strategies across different model configurations remains challenging. In this work, we investigate the transferability of annealing dynamics in large language model training and refine a generalized predictive framework for optimizing annealing strategies under the Warmup-Steady-Decay (WSD) scheduler. Our improved framework incorporates training steps, maximum learning rate, and annealing behavior, enabling more efficient optimization of learning rate schedules. Our work provides a practical guidance for selecting optimal annealing strategies without exhaustive hyperparameter searches, demonstrating that smaller models can serve as reliable proxies for optimizing the training dynamics of larger models. We validate our findings on extensive experiments using both Dense and Mixture-of-Experts (MoE) models, demonstrating that optimal annealing ratios follow consistent patterns and can be transferred across different training configurations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。