arXiv:2512.02342math.OCcs.LG2025-12被引 4

提出新梯度步长,让非光滑优化更稳定且无需小梯度

Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)Gradients

  • 设计安全版随机Polyak步长,不依赖最优解或插值假设
  • 在凸优化中收敛性有理论保障,深度网络训练也表现稳定
  • 对梯度消失有鲁棒性,适合难调参的非光滑场景

随机Polyak步长(SPS)在光滑凸与非凸优化问题中表现出色,尤其在深度神经网络训练中表现优异。然而,其在非光滑场景下的扩展仍处于初级阶段,常依赖插值假设或需知晓最优解。本文提出一种新的SPS变体——安全版SPS(SPS$_{safe}$),适用于随机次梯度法,并在无强假设条件下为非光滑凸优化提供了严格的收敛保证。我们进一步将动量机制引入更新规则,同样获得紧致的理论结果。大量实验验证了理论:该步长在凸基准测试和深度网络训练中均达到与现有自适应方法相当的性能,且在多种问题设置下表现稳定。更重要的是,在深度网络训练中,该步长下的梯度范数不会坍缩至接近零,表明对梯度消失具有鲁棒性。

原文摘要 · Abstract (English)

The stochastic Polyak step size (SPS) has proven to be a promising choice for stochastic gradient descent (SGD), delivering competitive performance relative to state-of-the-art methods on smooth convex and non-convex optimization problems, including deep neural network training. However, extensions of this approach to non-smooth settings remain in their early stages, often relying on interpolation assumptions or requiring knowledge of the optimal solution. In this work, we propose a novel SPS variant, Safeguarded SPS (SPS$_{safe}$), for the stochastic subgradient method, and provide rigorous convergence guarantees for non-smooth convex optimization with no need for strong assumptions. We further incorporate momentum into the update rule, yielding equally tight theoretical results. Comprehensive experiments on convex benchmarks and deep neural networks corroborate our theory: the proposed step size achieves competitive performance to existing adaptive baselines and exhibits stable behavior across a wide range of problem settings. Finally, in the context of deep neural network training, the gradient norms under our step size do not collapse to (near) zero, indicating robustness to vanishing gradients.

优化算法非光滑优化深度学习梯度消失

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。