arXiv:2606.15702math.OCcs.LG2026-06

提出基于施瓦茨范数的自适应优化方法,统一了经典与新型优化器。

Schattor: Schatten-family methods for deep learning optimization

  • 基于施瓦茨范数构建自适应优化框架,融合SGD与Muon优点。
  • 在矩阵优化中实现无维度依赖的稳定点保证,理论更可靠。
  • 适用于高维矩阵模型优化,尤其适合结构复杂深度学习任务。

现代深度学习优化面临参数结构异质、梯度噪声大、非凸性高等挑战,对算法设计与理论分析均构成难题。受SGD局限与自适应优化器成功启发,我们提出{ m Schattor}——一类基于施瓦茨范数的自适应一阶方法。Schattor将SGD与近期提出的矩阵型自适应优化器Muon统一于单一施瓦茨范数框架下。通过新颖的矩阵鞅矩界,我们为该家族中的方法建立了针对随机矩阵优化问题的无维度稳定点保证。此外,我们还发展了多块扩展版本,可自适应平衡各块优化进度,并在更一般设置下证明了无维度稳定点保证。

原文摘要 · Abstract (English)

Modern deep learning optimization features heterogeneous parameter structures, noisy gradients, and highly nonconvex landscapes, posing significant challenges for both algorithm design and theoretical analysis. Motivated by the limitations of SGD and the success of adaptive optimizers, we propose {\it Schattor}, a family of adaptive first-order methods based on Schatten norms. Schattor unifies SGD and the recently proposed matrix-variate adaptive optimizer Muon within a single Schatten-norm-based framework. We establish dimension-free stationarity guarantees for methods in the Schattor family for stochastic matrix optimization problems via a novel matrix martingale moment bound. We also develop multi-block extensions that adaptively balance block-wise optimization progress and prove dimension-free stationarity guarantees in this more general setting.

优化算法矩阵优化自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。