arXiv:2605.18609cs.LG2026-05

揭示了动量加速与小批量大小的正比关系,实现计算完全并行化。

Perfect Parallelization in Mini-Batch SGD with Classical Momentum Acceleration

  • 基于二次优化与插值假设,建立通用动量加速理论框架。
  • 动量加速效果与梯度小批量大小成正比,达饱和点前可完美并行。
  • 给出简单有效的动量参数选择,实证表现良好。

以经典动量方法(如Polyak动量)加速随机梯度方法在大规模机器学习训练中已取得显著成功,尤其结合大批次计算硬件加速时。然而,经典动量对随机小批量优化的影响仍缺乏充分的理论理解,以往研究依赖强噪声假设且要求极大数据批次。本文针对插值情形下的二次优化问题,建立了一般性理论框架,涵盖动量型与Nesterov型动量,适用于任意小批量大小,对随机噪声假设极少。我们证明:动量加速效果与梯度小批量大小成正比(至自然饱和点),从而实现小批量计算的完美并行化。该理论还提供了简单有效的动量参数选择,经实验证明有效。

原文摘要 · Abstract (English)

Accelerating stochastic gradient methods with classical momentum schemes, such as Polyak's heavy ball, has proven highly successful in training large-scale machine learning models, particularly when combined with the hardware acceleration of large mini-batch computations. Yet, the effect of classical momentum on stochastic mini-batch optimization has been poorly understood theoretically, with prior works requiring strong noise assumptions and extremely large mini-batches. In this work, we develop a general theory of stochastic momentum acceleration for optimizing over quadratics in the interpolation regime, a popular abstraction for studying deep learning dynamics which also includes classical methods such as randomized Kaczmarz and coordinate descent. Our framework encompasses both heavy ball and Nesterov-style momentum, allows for arbitrary mini-batch sizes, and makes minimal assumptions on the stochastic noise. In particular, we show that acceleration from classical momentum is directly proportional to the gradient mini-batch size (up to a natural saturation point), thereby enabling perfect parallelization of mini-batch computations. Our theory also provides a simple choice for the momentum parameter, which is shown to be effective empirically.

动量优化小批量并行计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。