arXiv:2506.14806cs.LG2025-06NeurIPS

提出可控制离散误差的连续时间动量模型,揭示其在深度学习中的隐式正则化作用。

Heavy-Ball Momentum Method in Continuous Time and Discretization Error Analysis

  • 设计分段连续微分方程,显式引入校正项以控制步长的离散误差。
  • 实现对离散误差的任意阶控制,提升连续逼近的理论精度。
  • 揭示动量方法在对角线线性网络中的方向平滑性隐式正则化机制。

本文建立了离散Heavy-Ball(HB)动量法的连续时间近似模型,即分段连续微分方程,并显式分析了离散化误差。尽管连续微分方程研究是理解离散优化方法的重要途径,但动量项带来的离散误差与连续模型之间的差距仍未被系统填补。本工作聚焦于连续时间下的HB方法,特别关注离散误差,设计了一阶分段连续微分方程,通过添加若干补偿项来显式建模离散误差。结果提供了一个可任意阶控制步长误差的连续时间模型。作为应用,我们据此发现了一种新的方向平滑性隐式正则化,并研究了HB方法在对角线线性网络中的隐式偏差,说明该理论可用于深度学习。数值实验进一步验证了理论结论。

原文摘要 · Abstract (English)

This paper establishes a continuous time approximation, a piece-wise continuous differential equation, for the discrete Heavy-Ball (HB) momentum method with explicit discretization error. Investigating continuous differential equations has been a promising approach for studying the discrete optimization methods. Despite the crucial role of momentum in gradient-based optimization methods, the gap between the original discrete dynamics and the continuous time approximations due to the discretization error has not been comprehensively bridged yet. In this work, we study the HB momentum method in continuous time while putting more focus on the discretization error to provide additional theoretical tools to this area. In particular, we design a first-order piece-wise continuous differential equation, where we add a number of counter terms to account for the discretization error explicitly. As a result, we provide a continuous time model for the HB momentum method that allows the control of discretization error to arbitrary order of the step size. As an application, we leverage it to find a new implicit regularization of the directional smoothness and investigate the implicit bias of HB for diagonal linear networks, indicating how our results can be used in deep learning. Our theoretical findings are further supported by numerical experiments.

优化算法动量方法连续逼近深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。