arXiv:2605.12878math.OCcs.LG2026-05被引 1

提出新型自适应优化器Adam-SHANG,提升随机凸优化收敛性。

Adam-SHANG: A Convergent Adam-Type Method for Stochastic Smooth Convex Optimization

论文配图:Adam-SHANG: A Convergent Adam-Type Method for Stochastic Smooth Convex Optimization
图 1 · 摘自论文原文
  • 基于李雅普诺夫函数设计稳定预条件更新机制
  • 理论证明在可调步长下期望收敛,无需全局单调性约束
  • 实测性能优于Adam和AdamW,适用于深度学习任务

我们提出Adam-SHANG,一种基于李雅普诺夫引导的亚当型方法,通过更稳定的滞后预条件更新耦合动量、自适应预处理和曲率感知修正。针对随机光滑凸优化问题,在可接受步长条件下证明了期望收敛性,该条件可通过保守谱界始终满足,且不需对二阶矩序列施加全局单调性。为获得更实用的步长规则,引入可计算的迹比步长,其依据是局部坐标对齐条件。相同结构更新在非凸设置下也进行了测试,参数简化后表现良好。实验验证了预测的随机衰减行为,并在深度学习任务中表现出与Adam和AdamW相当甚至更优的训练性能。

原文摘要 · Abstract (English)

We propose Adam-SHANG, a Lyapunov-guided Adam-type method that couples momentum, adaptive preconditioning, and a curvature-aware correction through a more stable lagged-preconditioner update. For stochastic smooth convex optimization, we prove convergence in expectation under an admissible stepsize condition that can always be satisfied by a conservative spectral bound, without imposing global monotonicity on the second-moment sequence. To obtain a less conservative practical rule, we introduce a computable trace-ratio stepsize, motivated by a local coordinatewise alignment condition. The same structural update is also tested beyond the convex setting with simplified parameters. Experiments validate the predicted stochastic decay and show competitive training performance against Adam and AdamW on deep learning tasks.

优化算法自适应优化深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。