arXiv:2602.16167cs.LG2026-02被引 3

提出新优化器SpecMuon,解决科学机器学习中的梯度病态问题。

Muon with Spectral Guidance: Efficient Optimization for Scientific Machine Learning

  • 基于奇异值分解的分模式更新,动态调节步长
  • 收敛速度比Adam、Muon快,且在多个物理方程上更稳定
  • 适合求解带刚性约束的偏微分方程神经网络

物理信息神经网络与神经算子常因梯度病态、多尺度谱行为及物理约束带来的刚性导致优化困难。近期提出的Muon优化器通过在梯度的奇异向量基下进行正交更新,改善了几何条件。但其单位奇异值更新可能引发过激步长,缺乏显式稳定性保证。本文提出SpecMuon,将Muon的正交几何与模式感知的松弛标量辅助变量(RSAV)机制结合。通过将矩阵梯度分解为奇异模式,并沿主导谱方向分别应用RSAV更新,SpecMuon根据全局损失能量自适应调节步长,同时保持原方法的尺度平衡性。该方法将优化视为多模式梯度流,实现对刚性谱成分的合理控制。我们建立了严格的理论性质:修正的能量耗散律、辅助变量的正性和有界性,以及在Polyak-Lojasiewicz条件下全局线性收敛。数值实验表明,在一维Burgers方程和分数阶偏微分方程等基准问题上,SpecMuon相比Adam、AdamW和原始Muon具有更快收敛速度和更高稳定性。

原文摘要 · Abstract (English)

Physics-informed neural networks and neural operators often suffer from severe optimization difficulties caused by ill-conditioned gradients, multi-scale spectral behavior, and stiffness induced by physical constraints. Recently, the Muon optimizer has shown promise by performing orthogonalized updates in the singular-vector basis of the gradient, thereby improving geometric conditioning. However, its unit-singular-value updates may lead to overly aggressive steps and lack explicit stability guarantees when applied to physics-informed learning. In this work, we propose SpecMuon, a spectral-aware optimizer that integrates Muon's orthogonalized geometry with a mode-wise relaxed scalar auxiliary variable (RSAV) mechanism. By decomposing matrix-valued gradients into singular modes and applying RSAV updates individually along dominant spectral directions, SpecMuon adaptively regulates step sizes according to the global loss energy while preserving Muon's scale-balancing properties. This formulation interprets optimization as a multi-mode gradient flow and enables principled control of stiff spectral components. We establish rigorous theoretical properties of SpecMuon, including a modified energy dissipation law, positivity and boundedness of auxiliary variables, and global convergence with a linear rate under the Polyak-Lojasiewicz condition. Numerical experiments on physics-informed neural networks, DeepONets, and fractional PINN-DeepONets demonstrate that SpecMuon achieves faster convergence and improved stability compared with Adam, AdamW, and the original Muon optimizer on benchmark problems such as the one-dimensional Burgers equation and fractional partial differential equations.

优化器物理信息神经网络收敛

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。