提出梯度裁剪在重尾噪声下的收敛性理论,突破传统假设限制。
Convergence of Clipped-SGD for Convex $(L_0,L_1)$-Smooth Optimization with Heavy-Tailed Noise
- 基于$(L_0,L_1)$-光滑性与重尾噪声,首次建立剪裁梯度法的高概率收敛界。
- 收敛速率不依赖亚高斯假设,避免指数级放大因子,适用范围更广。
- 适用于大语言模型训练等存在异常梯度的场景,理论指导性强。
梯度裁剪是机器学习与深度学习中广泛应用的技术,能有效缓解大语言模型训练中常见的重尾噪声影响。此外,带有裁剪的一阶方法(如Clip-SGD)在$(L_0,L_1)$-光滑性假设下,相比普通SGD具有更强的收敛保证,该性质在许多深度学习任务中被观察到。然而,现有文献尚未完整解决在重尾噪声与$(L_0,L_1)$-光滑性双重假设下Clip-SGD的高概率收敛问题。本文首次建立了凸$(L_0,L_1)$-光滑优化中带重尾噪声的Clip-SGD的高概率收敛界。分析结果扩展了已有研究:既恢复了确定性情形与$ L_1 = 0 $时的随机设定下的已知界限,又避免了指数级放大的因子,且不依赖于严格的亚高斯噪声假设,显著拓展了梯度裁剪的应用边界。
原文摘要 · Abstract (English)
Gradient clipping is a widely used technique in Machine Learning and Deep Learning (DL), known for its effectiveness in mitigating the impact of heavy-tailed noise, which frequently arises in the training of large language models. Additionally, first-order methods with clipping, such as Clip-SGD, exhibit stronger convergence guarantees than SGD under the $(L_0,L_1)$-smoothness assumption, a property observed in many DL tasks. However, the high-probability convergence of Clip-SGD under both assumptions -- heavy-tailed noise and $(L_0,L_1)$-smoothness -- has not been fully addressed in the literature. In this paper, we bridge this critical gap by establishing the first high-probability convergence bounds for Clip-SGD applied to convex $(L_0,L_1)$-smooth optimization with heavy-tailed noise. Our analysis extends prior results by recovering known bounds for the deterministic case and the stochastic setting with $L_1 = 0$ as special cases. Notably, our rates avoid exponentially large factors and do not rely on restrictive sub-Gaussian noise assumptions, significantly broadening the applicability of gradient clipping.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。