动态调整批量大小可显著提升大模型训练效率
Towards joint scaling laws with optimal batch size schedules
- 从凸优化视角建立学习率与批量大小的联合动态模型
- 推导出任意学习率调度下的最优批量大小方案
- 在大模型训练中表现优于固定批量大小基线
现代深度学习通常在训练过程中保持批量大小不变,从而忽略了学习率与批量大小对训练动态的联合影响。本文从凸优化视角研究深度学习动态,推导出适用于一般优化器和模型架构的损失联合表征,该表征揭示了学习率与批量大小调度的协同关系。基于此,我们获得任意给定学习率调度下的闭式最优批量大小调度,并进一步建立联合缩放律,其性能持续优于静态批量大小基线,凸显了动态批量大小调度在大语言模型训练中的重要性。
原文摘要 · Abstract (English)
Modern deep learning typically keeps the batch size static throughout training, thus overlooking the joint effect of learning rate and batch size on the training dynamics. In this paper, we study the deep learning dynamics through the lens of convex optimization and derive a joint characterization of loss in terms of both schedules, applicable to general optimizers and model architectures. This characterization yields a closed-form optimal batch size schedule for any prescribed learning rate schedule, and further leads to joint scaling laws that consistently outperform static batch size baselines, highlighting the significance of dynamic batch size schedule in large language model training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。