arXiv:2512.09678math.OCcs.AI2025-12被引 3

提出新型矩阵优化算法,提升深度学习训练效果。

Ky Fan Norms and Beyond: Dual Norms and Combinations for Matrix Optimization

  • 基于凯·范范数对偶构造新优化器家族
  • F-Muon与S-Muon在多任务中性能优于或等同于Muon
  • 适合追求高效稳定训练的深度学习研究者

本文探讨了多种矩阵范数在权重矩阵函数优化中的应用,这是深度学习中的关键问题。超越支撑Muon更新的谱范数,我们利用凯·范范数的对偶性,引入了与Muon、ν-SAM和Dion密切相关的一系列线性最小化预言机(LMO)算法,统称为Fanion族。在LMO框架内,进一步构建了F-Fanion和S-Fanion族,其更新为Fanion与归一化SGD或SignSGD更新的凸组合。三类算法在广泛任务与设置下的大量实证研究显示,其中最具前景的F-Muon和S-Muon在多数情况下性能与Muon相当,且在合成光滑凸问题上表现更优。

原文摘要 · Abstract (English)

In this article, we explore the use of various matrix norms for optimizing functions of weight matrices, a crucial problem in deep learning. Moving beyond the spectral norm that underlies the Muon update, we leverage the duals of the Ky Fan norms to introduce the Fanion family of linear minimization oracle (LMO) algorithms, which are closely related to Muon, $ν$-SAM, and Dion. Staying inside the LMO, we construct the families of F-Fanions and S-Fanions, whose updates are convex combinations of the updates of Fanions and Normalized SGD or SignSGD, respectively. The most promising algorithms in these families are F-Muon and S-Muon. By conducting an extensive empirical study of all three algorithm families across a wide range of tasks and settings, we demonstrate that F-Muon and S-Muon consistently match Muon's performance, while outperforming Muon on a synthetic smooth convex problem.

矩阵优化深度学习优化算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。