arXiv:2602.16340cs.LGstat.ML2026-02被引 5

揭示动量优化器在光滑齐次网络中的隐式偏好机制

The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks

  • 将动量梯度下降视为衰减学习率下的近似最速下降轨迹
  • 证明不同优化器分别趋向于对应范数的margin最大化解
  • 实验验证优化器选择影响最终收敛点,对模型设计有指导意义

我们研究了基于动量的优化器在光滑齐次模型中的隐式偏差。研究表明,如Muon(谱范数)、MomentumGD(ℓ₂范数)和Signum(ℓ∞范数)等动量最速下降算法,在衰减学习率调度下是近似最速下降轨迹,证明这些算法具有朝向相应边际最大化问题的KKT点的偏差。我们将分析扩展至Adam(无稳定常数项),其最大化ℓ∞边际;以及Muon-Signum与Muon-Adam,它们最大化混合范数。实验结果支持理论,并表明所最大化的边际类型取决于优化器的选择。总体而言,我们的结果拓展了早期关于齐次模型中最速下降及线性模型中动量优化器的研究。

原文摘要 · Abstract (English)

We study the implicit bias of momentum-based optimizers on smooth homogeneous models. We show that \textit{momentum steepest descent} algorithms like Muon (spectral norm), MomentumGD ($\ell_2$ norm), and Signum ($\ell_\infty$ norm) are \textit{approximate} steepest descent trajectories under a decaying learning rate schedule, proving that these algorithms have a bias towards KKT points of the corresponding margin maximization problem. We extend the analysis to Adam (without the stability constant), which maximizes the $\ell_\infty$ margin, and to Muon-Signum and Muon-Adam, which maximize a hybrid norm. Our experiments corroborate the theory and show that the identity of the margin maximized depends on the choice of optimizer. Overall, our results extend earlier lines of work on steepest descent in homogeneous models and momentum-based optimizers in linear models.

优化器分析隐式偏差动量方法齐次网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。