arXiv:2509.22938cs.LG2025-09被引 2

SOAP优化器通过梯度白化视角揭示其优于Adam的机制

Understanding SOAP from the Perspective of Gradient Whitening

  • 从梯度白化角度分析SOAP,将其预条件矩阵视为二阶曲率近似
  • 实验证明SOAP与Shampoo收敛速度相当,最终损失无显著差异
  • 适合研究优化器理论机制或想理解SOAP优势的研究者

最近提出的预条件器特征基下的Adam+Shampoo(SOAP)在语言建模任务中表现出优于Adam和Shampoo的训练效率。本文从梯度白化视角分析Adam、Shampoo与SOAP,将它们的预条件器解释为白化矩阵的近似,该矩阵捕获了二阶曲率信息。我们进一步在克罗内克积假设下建立了理想化SOAP与Shampoo之间的理论等价性。为验证这些洞见,我们使用nanoGPT和灰度图像色彩化重现了语言建模实验。结果表明,SOAP的收敛速率与Shampoo相近,且在最终损失上未显著优于Adam和Shampoo,与理论等价性一致。

原文摘要 · Abstract (English)

Shampoo with Adam in the Preconditioner's eigenbasis (SOAP) has recently emerged as a promising optimization algorithm for neural network training, achieving superior training efficiency over both Adam and Shampoo in language modeling tasks. In this work, we analyze Adam, Shampoo, and SOAP from the perspective of gradient whitening, interpreting their preconditioners as approximations to the whitening matrix, which captures second-order curvature information. We further establish a theoretical equivalence between idealized versions of SOAP and Shampoo under the Kronecker product assumption. To empirically evaluate these insights, we reproduce the language modeling experiments using nanoGPT and grayscale image colorization. Our results show that SOAP exhibits similar convergence rate as Shampoo, and no significant advantage over both Adam and Shampoo in the final loss achieved, which aligns with their equivalence in theory.

优化器梯度白化神经网络训练SOAP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。