arXiv:2601.12238stat.MLcs.LG2026-01中稿 · ICML被引 1

证明动量SGD在非平稳优化中会因梯度延迟导致跟踪滞后。

On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization

  • 分解跟踪误差为瞬态、噪声和漂移三部分,揭示动量的代价
  • 当动量参数趋近1时,漂移放大效应使跟踪误差发散
  • 在动态变化场景下,普通SGD反而比加速版更稳定可靠

本文对随机梯度下降(SGD)及其动量变体(Polyak Heavy-Ball 和 Nesterov)在强凸与光滑条件下追踪时变最优解的能力进行了全面的理论分析。有限时间界揭示了跟踪误差的精细分解:包括瞬态误差、噪声诱导误差和漂移诱导误差。这一分解暴露了一个根本性权衡:尽管动量常被用作梯度平滑的启发式方法,但在分布漂移下,其会引入显式的漂移放大惩罚,且当动量参数 $β$ 趋近于 1 时,该惩罚趋于发散,导致系统性跟踪滞后。我们进一步在梯度变化约束下给出极小极大下界,证明这种动量引发的跟踪代价并非分析缺陷,而是信息论上的不可逾越障碍——在漂移主导的场景中,动量不可避免地表现更差,因为旧梯度的平均化导致系统性延迟。结果为动量在动态设置中的经验不稳定性提供了理论依据,并精确界定了普通SGD优于加速版本的适用区间。

原文摘要 · Abstract (English)

In this paper, we provide a comprehensive theoretical analysis of Stochastic Gradient Descent (SGD) and its momentum variants (Polyak Heavy-Ball and Nesterov) for tracking time-varying optima under strong convexity and smoothness. Our finite-time bounds reveal a sharp decomposition of tracking error into transient, noise-induced, and drift-induced components. This decomposition exposes a fundamental trade-off: while momentum is often used as a gradient-smoothing heuristic, under distribution shift it incurs an explicit drift-amplification penalty that diverges as the momentum parameter $β$ approaches 1, yielding systematic tracking lag. We complement these upper bounds with minimax lower bounds under gradient-variation constraints, proving this momentum-induced tracking penalty is not an analytical artifact but an information-theoretic barrier: in drift-dominated regimes, momentum is unavoidably worse because stale-gradient averaging forces systematic lag. Our results provide theoretical grounding for the empirical instability of momentum in dynamic settings and precisely delineate regime boundaries where vanilla SGD provably outperforms its accelerated counterparts.

优化理论动量方法非平稳优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。