arXiv:2509.25783stat.MLcs.LG2025-09被引 4

首次给出深度矩阵分解中极小值尖锐度的精确表达式,揭示平坦极小值的谱范数平衡规律。

Sharpness of Minima in Deep Matrix Factorization

  • 推导出损失函数海森矩阵最大特征值的精确公式
  • 证明谱范数乘积恒定是平坦极小值的充要条件
  • 揭示梯度训练中逃逸现象的首例实证,适合研究优化几何者阅读

理解非凸优化问题(如深度神经网络训练与深度矩阵分解)中极小值附近的损失景观几何结构,对解释基于梯度方法的隐式偏差至关重要。刻画该几何的核心量是损失函数海森矩阵的最大特征值。然而,由于在一般设定下缺乏该尖锐度度量的精确表达式,其确切作用长期模糊。本文首次给出了深度矩阵分解/深度线性神经网络训练问题中任意极小值处平方误差损失海森矩阵最大特征值的精确表达式,解决了Mulayoff & Michaeli(2020)提出的开放问题。该表达式揭示了深度矩阵分解损失景观的根本性质:各层左右中间因子的谱范数乘积恒定是平坦性的充分条件。尤其在深度-2矩阵分解和深度过参数化标量分解中,该条件既是必要也是充分的,意味着平坦极小值具有谱范数平衡性,尽管未必满足Frobenius范数平衡。为补充理论,我们首次对深度矩阵分解问题中梯度训练接近极小值时的逃逸现象进行了实证刻画。

原文摘要 · Abstract (English)

Understanding the geometry of the loss landscape near a minimum is key to explaining the implicit bias of gradient-based methods in non-convex optimization problems such as deep neural network training and deep matrix factorization. A central quantity to characterize this geometry is the maximum eigenvalue of the Hessian of the loss. Currently, its precise role has been obfuscated because no exact expressions for this sharpness measure were known in general settings. In this paper, we present the first exact expression for the maximum eigenvalue of the Hessian of the squared-error loss at any minimizer in deep matrix factorization/deep linear neural network training problems, resolving an open question posed by Mulayoff & Michaeli (2020). This expression reveals a fundamental property of the loss landscape in deep matrix factorization: Having a constant product of the spectral norms of the left and right intermediate factors across layers is a sufficient condition for flatness. Most notably, in both depth-$2$ matrix and deep overparameterized scalar factorization, we show that this condition is both necessary and sufficient for flatness, which implies that flat minima are spectral-norm balanced even though they are not necessarily Frobenius-norm balanced. To complement our theory, we provide the first empirical characterization of an escape phenomenon during gradient-based training near a minimizer of a deep matrix factorization problem.

矩阵分解优化几何梯度训练尖锐度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。