arXiv:2504.12601cs.LGmath.OC2025-04被引 1

放宽步长条件,证明非凸优化中SGD几乎必然收敛

Stochastic Gradient Descent in Non-Convex Problems: Asymptotic Convergence with Relaxed Step-Size via Stopping Time Methods

  • 用停止时间法分析SGD,放松步长与梯度要求
  • 步长满足∑ε_t=∞且∑ε_t^p<∞(p>2)时可收敛
  • 无需全局Lipschitz连续,适合实际训练场景

随机梯度下降(SGD)在机器学习中广泛应用。以往在递减步长设置下的收敛性分析通常依赖Robbins-Monro条件,但实践中常采用更广泛的步长策略,现有结果仍受限且依赖强假设。本文提出基于停止时间的新分析框架,实现了在更宽松步长条件和较弱假设下对非凸问题中SGD的渐近收敛性分析。在非凸设定下,我们证明当步长序列{ε_t}_{t≥1}满足∑_{t=1}^{+∞} ε_t = +∞且∑_{t=1}^{+∞} ε_t^p < +∞(p > 2)时,SGD迭代点几乎必然收敛。相比先前研究,本分析消除了损失函数全局Lipschitz连续性的假设,并降低了对高阶随机梯度矩有界的限制。在此几乎必然收敛基础上,进一步建立了L_2收敛性。这些显著放宽的假设使理论结果更具普适性,从而提升其在实际场景中的适用性。

原文摘要 · Abstract (English)

Stochastic Gradient Descent (SGD) is widely used in machine learning research. Previous convergence analyses of SGD under the vanishing step-size setting typically require Robbins-Monro conditions. However, in practice, a wider variety of step-size schemes are frequently employed, yet existing convergence results remain limited and often rely on strong assumptions. This paper bridges this gap by introducing a novel analytical framework based on a stopping-time method, enabling asymptotic convergence analysis of SGD under more relaxed step-size conditions and weaker assumptions. In the non-convex setting, we prove the almost sure convergence of SGD iterates for step-sizes $ \{ ε_t \}_{t \geq 1} $ satisfying $\sum_{t=1}^{+\infty} ε_t = +\infty$ and $\sum_{t=1}^{+\infty} ε_t^p < +\infty$ for some $p > 2$. Compared with previous studies, our analysis eliminates the global Lipschitz continuity assumption on the loss function and relaxes the boundedness requirements for higher-order moments of stochastic gradients. Building upon the almost sure convergence results, we further establish $L_2$ convergence. These significantly relaxed assumptions make our theoretical results more general, thereby enhancing their applicability in practical scenarios.

优化算法非凸优化随机梯度收敛分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。