arXiv:2509.07972cs.LGmath.OC2025-09被引 4

提出新平滑性假设,证明学习率预热能显著加速模型收敛

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence

  • 构建新型平滑性假设,解释学习率预热的理论优势
  • 在特定条件下,预热使梯度下降快 Θ(T) 倍收敛
  • 为大规模训练中的学习率策略提供理论依据

学习率预热是大规模深度神经网络训练中广泛使用且有效的技术。尽管实践效果显著,其理论优势尚未被充分理解。为弥合理论与实践的差距,本文首次提出一类新的广义平滑性假设,并从理论和实证两方面验证其适用性。在此新假设下,研究了梯度下降(GD)在确定性和随机设置下的收敛性质。结果表明,学习率预热能持续加速GD收敛;在某些特定情形下,采用预热的GD比非递减学习率调度快最多Θ(T)倍,从优化理论角度揭示了该策略的优势。

原文摘要 · Abstract (English)

Learning rate warmup is a popular and practical technique in training large-scale deep neural networks. Despite the huge success in practice, the theoretical advantages of this strategy of gradually increasing the learning rate at the beginning of the training process have not been fully understood. To resolve this gap between theory and practice, we first propose a novel family of generalized smoothness assumptions, and validate its applicability both theoretically and empirically. Under the novel smoothness assumption, we study the convergence properties of gradient descent (GD) in both deterministic and stochastic settings. It is shown that learning rate warmup consistently accelerates GD, and GD with warmup can converge at most $Θ(T)$ times faster than with a non-increasing learning rate schedule in some specific cases, providing insights into the benefits of this strategy from an optimization theory perspective.

优化理论学习率收敛加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。