arXiv:2606.15268cs.LG2026-06

不同场景下该用哪种矩阵范数优化,关键看数据维度高低。

When to use what Schatten-$p$ norm in deep learning?

  • 根据低维场景需求,小阶Schatten范数更优
  • 提出适用于p>2的抗噪加速新理论,解释性能差异
  • 适合研究优化器设计或大规模训练的工程师

基于Schatten-∞范数的优化器(如Muon)表现出良好实际效果,但其是否有益存在看似矛盾的观察。我们通过分析发现,结论取决于具体使用场景。即使目标函数在Schatten-∞几何中光滑,当处于低维情形时(包括Chinchilla缩放规律),较小的Schatten-p几何仍可能更优。这一结论源于对SODA框架在p>2条件下新的抗噪声加速结果。该分析同时解释了Muon类方法为何无需预热、天然偏好大批次,并推导出任意p值下的批量大小缩放规则。

原文摘要 · Abstract (English)

Schatten-$\infty$ based optimizers such as Muon have shown promising empirical performance, but there remains seemingly conflicting observations regarding whether they are beneficial. We resolve this conflict by showing that the conclusion is regime dependent. Even when the objective is smooth in the Schatten-$\infty$ geometry, smaller Schatten-$p$ geometries can be optimal, specifically in the low-dimensional regime, which we show includes Chinchilla scaling. This conclusion follows from a new noise-robust acceleration result for the SODA framework for $p>2$. The same analysis explains why Muon-like methods do not require warmup, why they naturally favor large batches, and yields a batch size scaling rule for arbitrary $p$.

优化器矩阵范数深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。