arXiv:2605.15522math.OCcs.LG2026-05被引 1

提出更真实的梯度约束模型,证明剪裁版AdamW在非光滑优化中表现最优。

Stochastic Non-Smooth Convex Optimization with Unbounded Gradients

  • 引入广义Lipschitz函数,允许梯度随优化差距线性增长
  • 剪裁版AdamW在随机凸优化中实现最优全局收敛率
  • 适用于非光滑、无界梯度场景,适合深度学习等实际问题

现有非光滑优化理论多依赖梯度统一有界的强假设。本文提出一类更现实的广义Lipschitz函数类,其梯度范数由优化差距的仿射函数控制。针对此类问题,研究发现:剪裁版AdamW在随机优化中可严格优于SGD、AdaGrad等方法。分析揭示了AdamW中指数加权梯度累积的关键作用,而非简单平均。进一步证明剪裁版AdamW具有普适性,在广义光滑假设下获得更优收敛速率,并分析了对角与矩阵预条件下的收敛性,扩展至准凸设置。

原文摘要 · Abstract (English)

Much of the existing theory on first-order non-smooth optimization is built on a restrictive assumption that the gradients of the objective function are uniformly bounded. We introduce a much more realistic class of generalized Lipschitz functions, where the gradient norms are bounded by an affine function of the optimality gap. We then ask a natural question: what algorithm achieves the best global convergence rates for solving convex stochastic generalized Lipschitz optimization problems? To address this, we develop a new convergence analysis for several existing algorithms and find that AdamW with clipped updates, provably outperforms other popular stochastic optimization methods, such as SGD and AdaGrad. Moreover, our analysis establishes the critical role of AdamW's exponentially weighted gradient accumulation, as opposed to simple averaging. We further show that clipped AdamW is universal and achieves improved rates under the popular generalized smoothness assumption, analyze the convergence of clipped AdamW with diagonal and matrix preconditioners, and extend our results to the quasar-convex setting.

非光滑优化AdamW随机优化收敛分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。