arXiv:2601.20399math.OCcs.LG2026-01被引 1

提出新型随机子空间优化算法,提升重尾噪声下的收敛稳定性。

Convergence Analysis of Randomized Subspace Normalized SGD under Heavy-Tailed Noise

  • 在子空间中引入方向归一化,增强梯度更新鲁棒性。
  • 在重尾噪声下实现高概率收敛,理论复杂度优于全维归一化方法。
  • 适合大规模非凸优化场景,如深度学习中的不稳定梯度问题。

随机子空间方法可降低每轮计算开销;然而在非凸优化中,多数分析基于期望值,即使在次高斯噪声下,高概率界也极为稀缺。我们首先证明,随机子空间 SGD(RS-SGD)在次高斯噪声下具有高概率收敛界,其预言机复杂度与以往期望结果同阶。鉴于现代机器学习中重尾梯度普遍存在,我们进一步提出随机子空间归一化 SGD(RS-NSGD),将方向归一化融入子空间更新。假设噪声具有有界 $p$-th 幂矩,我们建立了期望与高概率收敛保证,并证明 RS-NSGD 可实现比全维归一化 SGD 更优的预言机复杂度。

原文摘要 · Abstract (English)

Randomized subspace methods reduce per-iteration cost; however, in nonconvex optimization, most analyses are expectation-based, and high-probability bounds remain scarce even under sub-Gaussian noise. We first prove that randomized subspace SGD (RS-SGD) admits a high-probability convergence bound under sub-Gaussian noise, achieving the same order of oracle complexity as prior in-expectation results. Motivated by the prevalence of heavy-tailed gradients in modern machine learning, we then propose randomized subspace normalized SGD (RS-NSGD), which integrates direction normalization into subspace updates. Assuming the noise has bounded $p$-th moments, we establish both in-expectation and high-probability convergence guarantees, and show that RS-NSGD can achieve better oracle complexity than full-dimensional normalized SGD.

优化算法重尾噪声子空间法收敛分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。