arXiv:2410.16561cs.LGmath.OC2024-10JMLR被引 17

重尾噪声下,梯度归一化可替代甚至优于梯度截断。

Revisiting Gradient Normalization and Clipping for Nonconvex SGD under Heavy-Tailed Noise: Necessity, Sufficiency, and Acceleration

  • 证明在光滑条件下,仅用归一化即可保证非凸SGD收敛。
  • 归一化与截断结合,在恶劣噪声下收敛速度显著提升。
  • 适合研究非凸优化中鲁棒性算法的设计者参考。

梯度截断长期以来被视为在重尾梯度噪声下保证随机梯度下降(SGD)收敛的关键手段。本文重新审视这一观点,探讨梯度归一化是否可作为有效替代或补充。我们证明,在个体光滑性假设下,仅使用梯度归一化即可确保非凸SGD的收敛性。此外,当与截断结合时,其在更严苛的噪声分布下能实现更快的收敛速率。本文提出统一理论,涵盖仅归一化、仅截断及二者结合三种策略。进一步分析现有方差缩减算法发现,在此设定下,仅归一化即足以保证收敛。最后,我们提出一种加速变体,在二阶光滑条件下可提升收敛性能。结果为重尾噪声下的非凸优化提供了理论依据与实践指导。

原文摘要 · Abstract (English)

Gradient clipping has long been considered essential for ensuring the convergence of Stochastic Gradient Descent (SGD) in the presence of heavy-tailed gradient noise. In this paper, we revisit this belief and explore whether gradient normalization can serve as an effective alternative or complement. We prove that, under individual smoothness assumptions, gradient normalization alone is sufficient to guarantee convergence of the nonconvex SGD. Moreover, when combined with clipping, it yields far better rates of convergence under more challenging noise distributions. We provide a unifying theory describing normalization-only, clipping-only, and combined approaches. Moving forward, we investigate existing variance-reduced algorithms, establishing that, in such a setting, normalization alone is sufficient for convergence. Finally, we present an accelerated variant that under second-order smoothness improves convergence. Our results provide theoretical insights and practical guidance for using normalization and clipping in nonconvex optimization with heavy-tailed noise.

非凸优化梯度归一化重尾噪声SGD

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。