arXiv:2512.23178math.OCcs.LG2025-12中稿 · ICLR被引 4

改进梯度裁剪在重尾噪声下的优化性能,理论更优且逼近最优。

Clipped Gradient Methods for Nonsmooth Convex Optimization under Heavy-Tailed Noise: A Refined Analysis

  • 引入广义有效维度,精细化分析梯度裁剪误差与概率不等式
  • 新速率比已有结果更快,强凸问题下突破现有理论上限
  • 首次证明期望收敛率的紧致性,适用于高鲁棒性机器学习场景

重尾噪声下的优化近年受到关注,因其更贴合现代机器学习任务的实际观测。传统假设梯度噪声具有有限二阶矩已被更现实的有界 𝔭 阶矩(𝔭∈(1,2])取代,记为上界 σₗᵖ。梯度裁剪是一种简单有效的应对方法。本文对裁剪随机梯度下降(Clipped SGD)进行精细化分析,提出两个新收敛速率:高概率情形下为 𝒪(σₗ dₑff⁻¹⁄²ᵖ ln¹⁻¹⁄ᵖ(1/δ) T¹⁄ᵖ⁻¹),强凸问题下为 𝒪(σₗ² dₑff⁻¹⁄ᵖ ln²⁻²⁄ᵖ(1/δ) T²⁄ᵖ⁻²),优于先前最优结果。其中,dₑff≥1 为提出的广义有效维度。分析改进体现在更优利用Freedman不等式和更精细的裁剪误差界。此外,将分析扩展至期望收敛,获得打破已知下界的速率。最后,建立了新的高概率与期望收敛下界,发现期望情形下新上界与下界匹配,表明分析在期望意义下达到最优。

原文摘要 · Abstract (English)

Optimization under heavy-tailed noise has become popular recently, since it better fits many modern machine learning tasks, as captured by empirical observations. Concretely, instead of a finite second moment on gradient noise, a bounded ${\frak p}$-th moment where ${\frak p}\in(1,2]$ has been recognized to be more realistic (say being upper bounded by $σ_{\frak l}^{\frak p}$ for some $σ_{\frak l}\ge0$). A simple yet effective operation, gradient clipping, is known to handle this new challenge successfully. Specifically, Clipped Stochastic Gradient Descent (Clipped SGD) guarantees a high-probability rate ${\cal O}(σ_{\frak l}\ln(1/δ)T^{1/{\frak p}-1})$ (resp. ${\cal O}(σ_{\frak l}^2\ln^2(1/δ)T^{2/{\frak p}-2})$) for nonsmooth convex (resp. strongly convex) problems, where $δ\in(0,1]$ is the failure probability and $T\in\mathbb{N}$ is the time horizon. In this work, we provide a refined analysis for Clipped SGD and offer two rates, ${\cal O}(σ_{\frak l}d_{\rm eff}^{-1/2{\frak p}}\ln^{1-1/{\frak p}}(1/δ)T^{1/{\frak p}-1})$ and ${\cal O}(σ_{\frak l}^2d_{\rm eff}^{-1/{\frak p}}\ln^{2-2/{\frak p}}(1/δ)T^{2/{\frak p}-2})$, faster than the aforementioned best results, where $d_{\rm eff}\ge1$ is a quantity we call the $\textit{generalized effective dimension}$. Our analysis improves upon the existing approach on two sides: better utilization of Freedman's inequality and finer bounds for clipping error under heavy-tailed noise. In addition, we extend the refined analysis to convergence in expectation and obtain new rates that break the known lower bounds. Lastly, to complement the study, we establish new lower bounds for both high-probability and in-expectation convergence. Notably, the in-expectation lower bounds match our new upper bounds, indicating the optimality of our refined analysis for convergence in expectation.

优化算法重尾噪声梯度裁剪理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。