提出新型稳定方法,让非凸优化学习更高效且更接近原始更新。
RELTA-SGLD: Relative-Growth Localized Taming for Nonconvex Stochastic-Gradient Langevin Learning
- 基于相对增长原则与阈值控制,动态调节梯度更新强度。
- 在Fashion-MNIST上优于未稳定SGLD和TUSLA,接近AdamW性能。
- 减少对原始更新的过度抑制,保持近无扰动的学习动态。
我们提出RELTA-SGLD,一种稳定超线性随机梯度更新的调制方案,同时降低对原始学习漂移的过度抑制。通过阈值确定调制启动位置,基于一步李雅普诺夫稳定性条件推导的相对增长原则决定所需调制强度。二者结合使λ尺度分母更轻,保留非消失的远尾返回。由此证明,在非凸SGLD中使用超线性增长的随机梯度算子时,可实现多项式矩稳定性及一阶驻点精度($W_1$ 和 $W_2$),优于同类调制方案的半阶与四分之一阶界。在时尚MNIST上主动稳定压力下,RELTA在均值学习指标上优于未调制SGLD与TUSLA,且与调参后的AdamW相当。在常规训练中,其更轻的局部化分母减少对原始更新的不必要扰动,维持近未调制的学习动态。
原文摘要 · Abstract (English)
We introduce RELTA-SGLD, a taming scheme that stabilizes superlinear stochastic-gradient updates while reducing unnecessary suppression of the original learning drift. A threshold determines where the taming turns on, while a relative-growth principle derived from the one-step Lyapunov stability condition determines the required taming strength. Together, they produce a lighter $λ$-scale denominator and preserve a nonvanishing far-tail return. As a consequence, we prove polynomial moment stability and first-order stationary accuracy in both $W_1$ and $W_2$ for nonconvex SGLD with superlinearly growing stochastic-gradient oracles, improving the corresponding half-order and quarter-order bounds for comparable stochastic-gradient tamed schemes. On Fashion-MNIST under active stabilization pressure, RELTA improves the mean learning metrics over both untamed SGLD and TUSLA and remains competitive with a tuned AdamW reference. In an ordinary-training regime, its lighter localized denominator reduces unnecessary perturbation of the original update and maintains nearly untamed learning dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。