提出OptMuon,让优化器自动调节更新幅度,提升稳定性与适应性。
OptMuon: Closed-Loop Orthogonalized Momentum Methods for Stochastic Optimization with Zero-Noise Optimality
- 采用轨迹依赖的自适应系数,根据实际梯度历史动态调整更新强度。
- 在零噪声下达到近最优的确定性收敛率,无需人工调参。
- 不依赖平滑常数或方差信息,对梯度突变有鲁棒性,适合高噪声场景。
正交动量更新在大规模深度学习中表现出强稳定性,但现有方法多依赖外部预设的固定或开环幅度规则,无法根据实际优化轨迹自适应调整。受无光滑性、自适应噪声方法的闭环思想启发,本文提出OptMuon,一类用于随机非凸优化的自适应正交动量方法。OptMuon将Muon风格的极坐标方向与依赖轨迹的AdaGrad-Norm型系数调度结合,使更新幅度由观测到的梯度和动量历史决定,而非依赖预设的光滑性常数。该调度不使用平滑常数、方差水平或有界梯度常数进行参数选择,其运行最大值修正可防止孤立梯度尖峰导致系数过度衰减。在下有界、无偏随机梯度具有有界方差、光滑性及几乎必然有界随机梯度条件下,证明了两个互补的期望平稳性保证:OptMuon-A在平均光滑性下达到噪声自适应率\(\tilde{\mathcal O}(T^{-1/2}+σ^{1/2}T^{-1/4})\),OptMuon-I在个体光滑性下达到\(\tilde{\mathcal O}(T^{-1/2}+σ^{1/3}T^{-1/3})\)。在零噪声情形下,两者均自动退化为近乎最优的确定性一阶率\(\tilde{\mathcal O}(T^{-1/2})\),无需手动重调超参数。结果表明,闭环标量自适应可与Muon风格正交动量结合,在保持噪声适应性的同时实现零噪声最优性(至对数因子)。
原文摘要 · Abstract (English)
Orthogonalized momentum updates, as used in Muon-style optimizers, have recently shown strong empirical stability in large-scale deep learning. However, most current orthogonalized methods are still paired with fixed, externally scheduled, or otherwise open-loop magnitude rules, so their scale is not directly calibrated from the realized optimization trajectory. Motivated by the closed-loop perspective behind Lipschitz-free and noise-adaptive methods, we propose OptMuon, a family of adaptive momentum orthogonalization methods for stochastic nonconvex optimization. OptMuon combines Muon-style polar-factor directions with a trajectory-dependent AdaGrad-Norm-type coefficient schedule, so that the update magnitude is determined by the observed gradient and momentum history rather than by a prescribed Lipschitz-dependent rule. The schedule does not use the smoothness constant, the variance level, or the bounded-gradient constant in parameter selection, and its running-maximum correction prevents isolated gradient spikes from causing excessive coefficient collapse. Under lower-boundedness, unbiased stochastic gradients with bounded variance, smoothness, and an almost-sure bounded stochastic-gradient condition, we prove two complementary expected-stationarity guarantees. OptMuon-A achieves the noise-adaptive rate \(\tilde{\mathcal O}(T^{-1/2}+σ^{1/2}T^{-1/4})\) under average smoothness, while OptMuon-I achieves \(\tilde{\mathcal O}(T^{-1/2}+σ^{1/3}T^{-1/3})\) under individual smoothness. In the zero-noise regime, both bounds automatically reduce to a nearly optimal deterministic first-order rate \(\tilde{\mathcal O}(T^{-1/2})\) without manual hyperparameter retuning. These results show that closed-loop scalar adaptation can be combined with Muon-style momentum orthogonalization while retaining noise adaptivity and zero-noise optimality up to logarithmic factors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。