arXiv:2605.09238cs.LGcs.AI2026-05被引 7

提出内在正则化方法,让矩阵优化在流形上更稳定高效

Intrinsic Muon: Spectral Optimization on Riemannian Matrix Manifolds

  • 基于黎曼流形的内在范数设计,保持对称性并支持闭式解
  • 在低秩、SPD、Stiefel等流形上实现通用闭式更新,收敛率仅依赖流形维度
  • 适用于大模型微调、图像分类等任务,无需重缩放参数

Muon及其相关范数约束的矩阵优化器已成为大规模学习问题的核心。它们被表述为在无约束欧氏空间中的矩阵范数球上的线性最大化查询(LMO)。然而,这些方法难以直接推广到流形值参数,如低秩分解、正交约束或对称正定(SPD)矩阵。将Muon LMO简单限制在切空间会破坏商对称性,并将切空间约束与环境范数界耦合,阻碍了多种感兴趣流形上的闭式解。我们通过一个关键观察解决了这两个问题:每个黎曼度量可自然地将酉不变欧氏范数提升为各切空间上的内在范数,由此产生的内在范数约束的LMO具有对称性保持性。基于此,我们提出了内在Muon(iMuon),一种统一框架,在任意酉不变范数(包括谱范数、Frobenius范数和核范数)下,对固定秩、SPD、Stiefel和格拉斯曼流形均能实现闭式更新。我们建立了确定性和随机iMuon的收敛保证,其速率常数仅依赖于流形维度。特别地,在固定秩流形上,该常数仅依赖于秩,使速率独立于因子条件,避免了先前工作所需的运行时因子重缩放。在LLM的LoRA微调、图像分类和子空间学习上的实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Muon and related norm-constrained matrix optimizers have become central to large-scale learning problems. They are formulated as a linear maximization oracle (LMO) over an ambient matrix-norm ball in unconstrained Euclidean space. However, these do not generalize cleanly to manifold-valued parameters such as low-rank factorizations, orthogonality constraints, or symmetric positive definite (SPD) matrices. Naively restricting the Muon LMO to the tangent space (i) breaks quotient symmetries and (ii) couples the tangent-space constraint with an ambient norm bound, thereby obstructing closed-form solutions on various manifolds of interest. We resolve both issues with a single observation: every Riemannian metric canonically lifts a unitarily invariant Euclidean norm to an intrinsic norm on each tangent space, and the resulting intrinsic norm constrained LMO is symmetry preserving. Building on this, we introduce intrinsic Muon (iMuon), a unified framework that yields closed-form updates on the fixed-rank, SPD, Stiefel, and Grassmann manifolds for any unitarily invariant norm, including the spectral, Frobenius, and nuclear norms. We establish convergence guarantees for both deterministic and stochastic iMuon with rate constants that depend only on the manifold dimension. Notably, on the fixed-rank manifold this constant depends only on the rank, making the rate independent of factor conditioning and removing the runtime factor-rescaling required by prior work. Experiments on LoRA finetuning of LLMs, image classification, and subspace learning illustrate the efficacy of the proposed approach.

矩阵优化黎曼流形大模型微调闭式解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。