提出新条件分析SGD收敛性,更贴近实际学习场景。
Convergence Analysis of SGD under Expected Smoothness
- 引入期望光滑性条件,放宽传统假设限制。
- 证明了多种步长策略下O(1/K)的收敛速率。
- 适合研究优化算法理论的学者参考。
随机梯度下降(SGD)是大规模学习的核心方法,但经典分析依赖过强或过粗的假设(如方差有界或噪声均匀)。期望光滑性(ES)条件作为灵活替代,将随机梯度的二阶矩与目标函数值及完整梯度关联。本文给出了在ES条件下的SGD自包含收敛性分析:(i)细化ES条件,引入可解释的采样依赖常数;(ii)推导完整梯度范数平方期望的上界;(iii)对多种步长策略证明带显式余项误差的O(1/K)收敛率。所有证明均在附录中详述。本工作统一并拓展了近期研究(Khaled and Richtárik, 2020;Umeda and Iiduka, 2025)。
原文摘要 · Abstract (English)
Stochastic gradient descent (SGD) is the workhorse of large-scale learning, yet classical analyses rely on assumptions that can be either too strong (bounded variance) or too coarse (uniform noise). The expected smoothness (ES) condition has emerged as a flexible alternative that ties the second moment of stochastic gradients to the objective value and the full gradient. This paper presents a self-contained convergence analysis of SGD under ES. We (i) refine ES with interpretations and sampling-dependent constants; (ii) derive bounds of the expectation of squared full gradient norm; and (iii) prove $O(1/K)$ rates with explicit residual errors for various step-size schedules. All proofs are given in full detail in the appendix. Our treatment unifies and extends recent threads (Khaled and Richtárik, 2020; Umeda and Iiduka, 2025).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。