研究带重尾噪声的非凸优化泛化能力,提出新分析框架
Stability and Generalization of Nonconvex Optimization with Heavy-Tailed Noise
- 通过截断技巧,基于有界p阶中心矩建立泛化界
- 证明了剪裁/归一化SGD及其变体在重尾噪声下仍稳定泛化
- 适用于训练过程存在异常梯度的现实场景
实证表明,使用重尾梯度噪声比传统有界方差噪声更适合作为机器学习训练过程的建模方式。现有研究多关注优化误差收敛性,而对重尾噪声下的泛化界分析仍不充分。本文提出一个通用框架,通过截断方法,在假设梯度具有有界p阶中心矩(p∈(1,2])条件下,推导出算法稳定性与泛化误差界的关联。基于此框架,进一步分析了多种主流随机算法在重尾噪声下的稳定性与泛化性能,包括剪裁与归一化随机梯度下降,及其小批量和动量变体。
原文摘要 · Abstract (English)
The empirical evidence indicates that stochastic optimization with heavy-tailed gradient noise is more appropriate to characterize the training of machine learning models than that with standard bounded gradient variance noise. Most existing works on this phenomenon focus on the convergence of optimization errors, while the analysis for generalization bounds under the heavy-tailed gradient noise remains limited. In this paper, we develop a general framework for establishing generalization bounds under heavy-tailed noise. Specifically, we introduce a truncation argument to achieve the generalization error bound based on the algorithmic stability under the assumption of bounded $p$th centered moment with $p\in(1,2]$. Building on this framework, we further provide the stability and generalization analysis for several popular stochastic algorithms under heavy-tailed noise, including clipped and normalized stochastic gradient descent, as well as their mini-batch and momentum variants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。