arXiv:2603.26554cs.LGstat.ML2026-03被引 7

揭示谱优化器在记忆存储中的高效性,解释为何其性能优于传统方法。

Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory

  • 通过高斯输入建模,突破正交嵌入限制,提升记忆容量。
  • Muon 存储容量远超 SGD,逼近牛顿法且仅用一阶信息。
  • 适合研究优化器机制或设计高效训练算法的研究者。

谱优化器如 Muon 在大规模语言模型训练中表现出色,但其优势来源和程度尚不明确。本文通过线性关联记忆问题——一种变压器模型事实回忆的可处理模型——进行研究。特别地,我们超越正交嵌入,考虑高斯输入与输出,使得存储关联数远超嵌入维度。主要结果精确刻画了在幂律频率分布下,Muon、SGD 与牛顿法在逻辑回归损失上单步恢复率。结果显示,Muon 的存储容量显著超过 SGD,甚至接近牛顿法,同时仅使用一阶信息。此外,Muon 在更大临界批量下饱和。进一步分析阈值梯度近似下的多步动态表明,Muon 初期恢复速度远快于 SGD,但两者最终以相近速度收敛至信息论极限。合成任务实验验证了预测的标度律。本分析为谱预条件器的信号放大提供了量化理解,并为建立更实际语言建模任务与优化器间的标度律奠定了基础。

原文摘要 · Abstract (English)

Spectral optimizers such as Muon have recently shown strong empirical performance in large-scale language model training, but the source and extent of their advantage remain poorly understood. We study this question through the linear associative memory problem, a tractable model for factual recall in transformer-based models. In particular, we go beyond orthogonal embeddings and consider Gaussian inputs and outputs, which allows the number of stored associations to greatly exceed the embedding dimension. Our main result sharply characterizes the recovery rates of one step of Muon, SGD, and Newton's method on the logistic regression loss under a power law frequency distribution. We show that the storage capacity of Muon significantly exceeds that of SGD, and even matches Newton's method while only using first-order information. Moreover, Muon saturates at a larger critical batch size. We further analyze the multi-step dynamics under a thresholded gradient approximation and show that Muon achieves a substantially faster initial recovery rate than SGD, while both methods eventually converge to the information-theoretic limit at comparable speeds. Experiments on synthetic tasks validate the predicted scaling laws. Our analysis provides a quantitative understanding of the signal amplification of spectral preconditioners and lays the groundwork for establishing scaling laws across more practical language modeling tasks and optimizers.

优化器记忆存储谱方法深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。