为高维随机梯度下降提供严格统计保证,解决理论空白
Statistical Guarantees for High-Dimensional Stochastic Gradient Descent
- 将高维时间序列工具引入在线学习,将SGD视为非线性自回归过程
- 证明常学习率下SGD在任意ℓ^s范数下的q阶矩收敛(q≥2)
- 给出高维ASGD的精确概率浓度界,适合高维稀疏建模研究者
随机梯度下降(SGD)及其Ruppert-Polyak平均变体(ASGD)是现代大规模学习的核心,但在高维情形下的理论性质仍不清晰。本文为常学习率的SGD和ASGD提供了严格的统计保证。关键创新在于将高维时间序列中的强大工具迁移至在线学习:将SGD视为非线性自回归过程,并采用现有耦合技术,证明了高维SGD在常学习率下的几何矩收缩,从而建立了迭代点的渐近平稳性。在此基础上,我们推导出SGD与ASGD在任意q≥2的ℓ^s-范数下的q阶矩收敛性,特别是广泛用于高维稀疏或结构化模型的ℓ^∞-范数。此外,我们提供了尖锐的高概率浓度分析,给出了高维ASGD的概率界。本工作不仅填补了SGD理论的关键空白,还为分析一类广泛的高维学习算法提供了新工具。
原文摘要 · Abstract (English)
Stochastic Gradient Descent (SGD) and its Ruppert-Polyak averaged variant (ASGD) lie at the heart of modern large-scale learning, yet their theoretical properties in high-dimensional settings are rarely understood. In this paper, we provide rigorous statistical guarantees for constant learning-rate SGD and ASGD in high-dimensional regimes. Our key innovation is to transfer powerful tools from high-dimensional time series to online learning. Specifically, by viewing SGD as a nonlinear autoregressive process and adapting existing coupling techniques, we prove the geometric-moment contraction of high-dimensional SGD for constant learning rates, thereby establishing asymptotic stationarity of the iterates. Building on this, we derive the $q$-th moment convergence of SGD and ASGD for any $q\ge2$ in general $\ell^s$-norms, and, in particular, the $\ell^{\infty}$-norm that is frequently adopted in high-dimensional sparse or structured models. Furthermore, we provide sharp high-probability concentration analysis which entails the probabilistic bound of high-dimensional ASGD. Beyond closing a critical gap in SGD theory, our proposed framework offers a novel toolkit for analyzing a broad class of high-dimensional learning algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。