arXiv:2410.16340stat.MLcs.LG2024-10被引 7

研究无穷方差梯度下的SGD收敛行为,揭示其渐近分布特性。

Limit Theorems for Stochastic Gradient Descent with Infinite Variance

  • 基于指数α∈(1,2)的正则变化梯度假设,扩展经典结果至多维情形。
  • 证明SGD渐近分布等价于稳定莱维过程驱动的欧尔恩斯坦-乌伦贝克过程的平稳分布。
  • 适用于高阶矩缺失的机器学习场景,如金融建模或极端事件预测。

随机梯度下降(SGD)是机器学习中训练模型的主流方法。尽管在有限方差梯度假设下已有深入研究,但对无穷方差梯度情形的理论分析仍不足。本文研究了当随机梯度服从指数α∈(1,2)的正则变化分布时,SGD的渐近行为。此前最接近的结果为1969年的一维情形,且限制于更窄的分布类。本文将其推广至多维情形,覆盖更广泛的无穷方差分布。我们发现,SGD的渐近分布可表征为由适当稳定莱维过程驱动的欧尔恩斯坦-乌伦贝克过程的平稳分布。此外,本文还探讨了这些结果在线性回归与逻辑回归模型中的应用。

原文摘要 · Abstract (English)

Stochastic gradient descent is a classic algorithm that has gained great popularity especially in the last decades as the most common approach for training models in machine learning. While the algorithm has been well-studied when stochastic gradients are assumed to have a finite variance, there is significantly less research addressing its theoretical properties in the case of infinite variance gradients. In this paper, we establish the asymptotic behavior of stochastic gradient descent in the context of infinite variance stochastic gradients, assuming that the stochastic gradient is regular varying with index $α\in(1,2)$. The closest result in this context was established in 1969 , in the one-dimensional case and assuming that stochastic gradients belong to a more restrictive class of distributions. We extend it to the multidimensional case, covering a broader class of infinite variance distributions. As we show, the asymptotic distribution of the stochastic gradient descent algorithm can be characterized as the stationary distribution of a suitably defined Ornstein-Uhlenbeck process driven by an appropriate stable Lévy process. Additionally, we explore the applications of these results in linear regression and logistic regression models.

优化理论随机梯度稳定分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。