揭示大模型训练中学习率与批量大小的最优动态规律。
Time Transfer: On Optimal Learning Rate and Batch Size In The Infinite Data Limit
- 发现最优学习率随预训练令牌数变化,且临界批量大小与令牌预算成正比。
- 最优批量大小会随时间增长而上升,固定批量会逐渐失效。
- 该规律在不同模型规模下仍成立,为联合数据与模型扩展提供新视角。
大规模语言模型(LLM)最优缩放面临超参数调优成本高昂的问题,尤其是学习率η和批量大小B。尽管μP方法(Yang et al., 2022)提供了无限模型规模下的最优η转移规则,但无限数据规模下的最优缩放行为仍未知。本文首次观察到最优η对预训练令牌预算T、批量大小B及其与临界批量大小B_crit的复杂依赖关系,测得B_crit ∝ T。同时发现最优批量大小与B_crit正相关:即使学习率被最优调整,固定批量也会随时间变得次优。令人意外的是,这些最优η和B的动态在μP模型缩放下依然保持,挑战了传统认为B_crit仅依赖损失值的观点。此外,我们发现损失对学习率变化的敏感性随T增加而降低,且在μP缩放下保持不变。本研究为数据与模型联合最优缩放的统一图景迈出第一步。
原文摘要 · Abstract (English)
One of the main challenges in optimal scaling of large language models (LLMs) is the prohibitive cost of hyperparameter tuning, particularly learning rate $η$ and batch size $B$. While techniques like $μ$P (Yang et al., 2022) provide scaling rules for optimal $η$ transfer in the infinite model size limit, the optimal scaling behavior in the infinite data size limit remains unknown. We fill in this gap by observing for the first time an intricate dependence of optimal $η$ scaling on the pretraining token budget $T$, $B$ and its relation to the critical batch size $B_\mathrm{crit}$, which we measure to evolve as $B_\mathrm{crit} \propto T$. Furthermore, we show that the optimal batch size is positively correlated with $B_\mathrm{crit}$: keeping it fixed becomes suboptimal over time even if learning rate is scaled optimally. Surprisingly, our results demonstrate that the observed optimal $η$ and $B$ dynamics are preserved with $μ$P model scaling, challenging the conventional view of $B_\mathrm{crit}$ dependence solely on loss value. Complementing optimality, we examine the sensitivity of loss to changes in learning rate, where we find the sensitivity to decrease with increase of $T$ and to remain constant with $μ$P model scaling. We hope our results make the first step towards a unified picture of the joint optimal data and model scaling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。