用谱沃斯泰因流解释深度学习参数更新的稳定机制
Muon Dynamics as a Spectral Wasserstein Flow
- 引入基于矩阵范数的谱沃斯泰因距离,统一多种归一化方法
- 证明在广义范数下,训练过程可视为梯度流,理论更统一
- 适用于宽网络、注意力模型等场景,适合研究优化稳定性者
梯度归一化能稳定深度学习优化,谱归一化对矩阵参数块尤为自然,以Muon为例。本文在均场框架下研究理想化的确定性连续时间无动量版本:宽网络由参数空间上的概率测度表示。从归一化矩阵流出发,定义依赖正半定矩阵范数γ的谱沃斯泰因距离:迹范数对应经典W₂,算子范数对应Muon几何,施加特范数则介于两者之间。我们建立了静态康托罗维奇形式化,提出最大-最小鲁棒代价表示,发展高斯简化并扩展布雷斯公式,并对单调范数证明其与贝纳穆-布伦耶形式等价。这为均场归一化训练动态提供了梯度流解释。数值实验验证了MMD流、高斯简化、两层ReLU模型及浅层注意力中的有效性。
原文摘要 · Abstract (English)
Gradient normalization stabilizes deep-learning optimization, and spectral normalizations are especially natural for matrix-shaped parameter blocks; Muon is the motivating example. We study an idealized deterministic, continuous-time, vanishing-momentum version of this idea in the mean-field regime, where wide models are represented by probability measures on parameter space. Starting from normalized matrix flows, we introduce Spectral Wasserstein distances indexed by norms $γ$ on positive semidefinite matrices: the trace norm gives classical $W_2$, the operator norm gives the Muon geometry, and Schatten norms interpolate between them. We develop the static Kantorovich formulation, a max-min robust-cost representation, Gaussian reductions extending the Bures formula, and for monotone norms, prove equivalence with a Benamou--Brenier formulation. This yields a gradient-flow interpretation of the mean-field normalized training dynamics. We illustrate these findings by numerical experiments on MMD flows, Gaussian reductions, two-layer ReLU models, and shallow attention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。