arXiv:2605.23871stat.MLcs.LG2026-05被引 1

从概率流角度重新解释Muon优化器,揭示其能量衰减机制。

Move on Muon : A Hamiltonian probability gradient flow perspective of Muon optimizer

论文配图:Move on Muon : A Hamiltonian probability gradient flow perspective of Muon optimizer
图 1 · 摘自论文原文
  • 将Muon更新视为概率空间中的镜像步,动量为对偶变量
  • 推导出带惯性的连续时间极限,能量单调下降
  • 适用于神经网络和Transformer模型的均场优化

我们构建了在矩阵参数概率测度空间上的梯度流,基于正则化Muon优化器——一种理想化Muon的解析平滑版本。核心发现是正则化正交化映射是核范数平滑Fenchel对偶的梯度。这表明(正则化)Muon更新可视为在更新变量上的镜像/近端步,动量作为对偶坐标。利用该结构,我们将Muon从单个矩阵参数推广至形式为 $J(ρ)=Rig(igint F d ρig)$ 的有限粒子概率目标,这一设定源于神经网络训练的均场描述,并推导出惯性连续时间极限。在此基础上,得到参数-动量对概率律上的相空间均场方程。所得到的流可被证明为阻尼哈密顿概率动力学,其动能由正则化Muon镜像势产生。我们证明了精确的哈密顿耗散恒等式,显示哈密顿能量单调递减。尽管目标函数本身未必沿惯性Muon动态单调,但在额外的梯度主导性、有界动量及曲率/对齐假设下,我们获得了目标间隙的连续与离散时间指数收敛率。同时研究了均场极限方程的适定性,并建立了交互粒子系统的混沌传播保证。最后,我们将该框架扩展至乘积矩阵空间上的希尔伯特值特征映射,得到了适用于平滑Transformer混合专家模型的分块Muon概率流。

原文摘要 · Abstract (English)

We develop a gradient flow on the space of probability measures defined on matrix-valued parameters induced by regularized Muon, an analytically smoothed version of the idealized Muon optimizer. The key observation is that the regularized orthogonalization map is the gradient of a smooth Fenchel-dual smoothing of the nuclear norm. This identifies the (regularized) Muon update as a mirror/prox step in the update variable, with momentum acting as the dual coordinate. We use this structure to lift Muon from a single matrix parameter to finite-particle probability objectives of the form $J(ρ)=R\left(\int F d ρ\right)$, a setting motivated by mean-field descriptions of neural-network training, and derive the inertial continuous-time limit. Using this structure, we derive the finite-particle continuous-time limit under the inertial scaling of step size and momentum, and then pass to a phase-space mean-field equation over probability laws on parameter-momentum pairs. The resulting flow can be shown to be a damped Hamiltonian probability dynamics whose kinetic energy is induced by the regularized Muon mirror potential. We prove an exact Hamiltonian dissipation identity, showing that the Hamiltonian energy decreases monotonically. While the target objective itself need not be monotone along the inertial Muon dynamics, under additional gradient-dominance, bounded-momentum, and curvature/alignment assumptions, we obtain continuous and discrete-time exponential convergence rates for the objective gap. We also study the well-posedness of the mean-field limit equation and establish propagation of chaos guarantees for the interacting particle system. Finally, we extend the formulation to Hilbert-valued feature maps on product matrix spaces, yielding a blockwise Muon probability flow applicable to smooth transformer mixture-of-experts models.

优化算法概率流神经网络哈密顿系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。