arXiv:2603.27072stat.MLcs.LG2026-03被引 1

证明正则化深度矩阵分解的解唯一且梯度平坦,为深层网络训练提供理论支撑。

On the Loss Landscape Geometry of Regularized Deep Matrix Factorization: Uniqueness and Sharpness

  • 通过L2正则化确保深度矩阵分解的全局唯一最优解
  • 所有极小值点的层范数相同,且海森矩阵迹有全局下界
  • 发现正则化参数临界阈值,超过后解退化为零

权重衰减在深层神经网络训练中广泛使用,其经验成功常归因于容量控制,但其对损失曲面结构和极小值集的影响仍缺乏理论理解。本文证明,在平方误差损失下,带ℓ²正则化的深度矩阵分解/深度线性网络问题,对任意目标矩阵,除一个勒贝格测度为零的深度与正则化参数集合外,均存在唯一的端到端极小值解。该结果揭示了正则化深度矩阵分解损失曲面的基本性质:所有极小值点的海森谱恒定。此外,若目标矩阵不在测度为零的集合中,则各层的弗罗比尼乌斯范数在所有极小值点上保持一致,并由此导出任意极小值点处海森矩阵迹的全局下界。进一步建立了正则化参数的临界阈值,当其超过该阈值时,唯一极小值解坍缩至零。

原文摘要 · Abstract (English)

Weight decay is ubiquitous in training deep neural network architectures. Its empirical success is often attributed to capacity control; nonetheless, our theoretical understanding of its effect on the loss landscape and the set of minimizers remains limited. In this paper, we show that $\ell^2$-regularized deep matrix factorization/deep linear network training problems with squared-error loss admit a unique end-to-end minimizer for all target matrices subject to factorization, except for a set of Lebesgue measure zero formed by the depth and the regularization parameter. This observation reveals fundamental properties of the loss landscape of regularized deep matrix factorization problems: the Hessian spectrum is constant across all minimizers of the regularized deep scalar factorization problem with squared-error loss. Moreover, we show that, in regularized deep matrix factorization problems with squared-error loss, if the target matrix does not belong to the Lebesgue measure-zero set, then the Frobenius norm of each layer is constant across all minimizers. This, in turn, yields a global lower bound on the trace of the Hessian evaluated at any minimizer of the regularized deep matrix factorization problem. Furthermore, we establish a critical threshold for the regularization parameter above which the unique end-to-end minimizer collapses to zero.

深度学习矩阵分解正则化优化几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。