给出平滑非线性神经网络损失曲率的闭式上界,无需数值计算即可评估泛化能力。
Wolkowicz-Styan Upper Bound on the Hessian Eigenspectrum for Cross-Entropy Loss in Nonlinear Smooth Neural Networks

- 基于Wolkowicz-Styan不等式推导交叉熵损失的海森矩阵最大特征值闭式上界
- 上界依赖于仿射参数、隐层维度和样本正交度,可直接反映模型复杂度影响
- 为理解深层网络泛化性提供新理论工具,适合研究损失几何与泛化关系的学者
神经网络是现代机器学习的核心,在诸多应用中取得顶尖性能。然而,损失函数几何结构与泛化能力之间的关系仍不清楚。在临界点附近,损失函数的局部几何由其二次型近似,该二次型来自二阶泰勒展开,其系数对应海森矩阵,其特征谱可用于衡量临界点处损失的尖锐程度。研究表明平坦的临界点泛化性能更好,而尖锐的临界点导致更高泛化误差。但海森特征谱的计算缺乏解析解,现有研究多依赖数值近似。已有闭式分析主要局限于线性或ReLU激活网络,对平滑非线性多层网络的理论分析仍有限。本文针对平滑非线性多层神经网络,利用Wolkowicz-Styan不等式,推导出交叉熵损失海森矩阵最大特征值的闭式上界。该上界表达为仿射变换参数、隐藏层维度及训练样本间正交度的函数。本文主要贡献在于通过闭式表达,首次对平滑非线性多层网络的损失尖锐性进行解析刻画,避免了显式的特征谱数值计算,为揭开深度学习之谜提供了有意义的一步。
原文摘要 · Abstract (English)
Neural networks (NNs) are central to modern machine learning and achieve state-of-the-art results in many applications. However, the relationship between loss geometry and generalization is still not well understood. The local geometry of the loss function near a critical point is well-approximated by its quadratic form, obtained through a second-order Taylor expansion. The coefficients of the quadratic term correspond to the Hessian matrix, whose eigenspectrum allows us to evaluate the sharpness of the loss at the critical point. Extensive research suggests flat critical points generalize better, while sharp ones lead to higher generalization error. However, sharpness requires the Hessian eigenspectrum, but general matrix characteristic equations have no closed-form solution. Therefore, most existing studies on evaluating loss sharpness rely on numerical approximation methods. Existing closed-form analyses of the eigenspectrum are primarily limited to simplified architectures, such as linear or ReLU-activated networks; consequently, theoretical analysis of smooth nonlinear multilayer neural networks remains limited. Against this background, this study focuses on nonlinear, smooth multilayer neural networks and derives a closed-form upper bound for the maximum eigenvalue of the Hessian with respect to the cross-entropy loss by leveraging the Wolkowicz-Styan bound. Specifically, the derived upper bound is expressed as a function of the affine transformation parameters, hidden layer dimensions, and the degree of orthogonality among the training samples. The primary contribution of this paper is an analytical characterization of loss sharpness in smooth nonlinear multilayer neural networks via a closed-form expression, avoiding explicit numerical eigenspectrum computation. We hope that this work provides a small yet meaningful step toward unraveling the mysteries of deep learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。