arXiv:2606.00542cs.LG2026-06

改进克罗内克分解优化器,通过选择不同散度提升训练效率

Rethinking Bregman Divergences in Kronecker-Factored Optimizers

  • 用贝格曼散度重新审视克罗内克分解的误差分布
  • 顶部特征空间与海森矩阵对齐度提高,尾部噪声更显著
  • 提出分层优化策略:顶层加权预条件,底层自适应加速

Shampoo 类优化器通过克罗内克分解结构近似梯度协方差矩阵。近期研究显示,此类近似可视为在贝格曼矩阵散度下的投影,从而生成不同的克罗内克预条件器。然而,当协方差并非精确克罗内克结构时,散度选择的作用仍不明确。本文通过协方差矩阵的谱分析研究该问题,发现弗罗贝尼乌斯、冯诺依曼和对数行列式散度会以不同方式分配不可避免的克罗内克近似误差。进一步表明,克罗内克因子由散度加权残差决定,而非原始误差,解释了其谱偏好如何体现在最终预条件器中。实验观察到,协方差的顶部特征空间与海森矩阵高度对齐,而尾部谱则明显噪声大且不可靠。基于此,我们提出一种子空间感知的克罗内克优化器,在顶部子空间使用基于特征值的预条件,在底部子空间采用自适应各向同性加速常数。

原文摘要 · Abstract (English)

Shampoo-style optimizers approximate gradient covariance matrices using Kronecker-factored structures. Recent work~\cite{lin2026understanding} showed that such approximations can be viewed as projections under Bregman matrix divergences, leading to different Kronecker-factored preconditioners. However, it remains unclear what role the choice of divergence plays when the covariance is not exactly Kronecker-factored. We study this question through the spectrum of the covariance matrix. We show that Frobenius, von Neumann, and LogDet divergences distribute the unavoidable Kronecker approximation error differently across the covariance spectrum. We further show that their Kronecker factors are governed by divergence-weighted residuals rather than the raw approximation error, explaining how these spectral preferences are realized in the resulting preconditioners. Empirically, we observe that the top covariance eigenspace is substantially better aligned with the Hessian matrix, while the tail spectrum is much noisier and unreliable. Motivated by these findings, we propose a subspace-aware Kronecker optimizer that applies eigenvalue-based preconditioning in the top subspace and uses an adaptive isotropic acceleration constant in the bottom subspace.

优化器克罗内克分解谱分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。