arXiv:2607.13631cs.LG2026-07中稿 · ICML

揭示神经网络海森谱与数据分布的关系,发现分类尖锐度取决于最大类别占比。

How the Hessian-Spectrum of Neural Networks Depends on Data

论文配图:How the Hessian-Spectrum of Neural Networks Depends on Data
图 1 · 摘自论文原文
  • 推导任意深度宽度线性网络的海森特征值,覆盖任意样本、特征和标签数。
  • 证明分类任务中解的尖锐度由最大类别样本比例决定,且该比例越高越尖锐。
  • 验证结论在非线性、实际数据假设下仍稳健,适合关注泛化与优化的研究者。

海森矩阵是研究深度学习损失曲面和优化动态的重要工具,也用于设计泛化度量和二阶优化算法。以往工作多依赖经验结果或在过于简化的理论假设下进行分析。本文推导了任意宽度与深度的线性网络在任意样本数、特征数和标签数下的海森特征值。特别地,在分类任务使用均方误差损失时,我们发现解的尖锐度直接与任一类样本的最大占比相关。我们通过实验验证了预测,并系统分析了逐个移除不切实际假设及引入非线性后的效果。结果表明,我们的预测在大多数情况下具有强鲁棒性,可推广至更实际的学习场景。

原文摘要 · Abstract (English)

The Hessian matrix is an important quantity of interest when it comes to studying the loss landscape and optimization dynamics in deep learning, as well as designing measures of generalization, second-order learning algorithms, etc. Prior works have focused on empirical results or pursued a theoretical treatment under overly simplified settings. In this work, we derive the eigenvalues of the Hessian of linear networks with arbitrary widths and depths, and datasets with an arbitrary number of samples, features, and labels. Importantly, for classification tasks with MSE loss, we identify that the sharpness of the solution is directly related to the maximum proportion of samples belonging to any class. We empirically validate our predictions and systematically analyze the effects of shedding the impractical assumptions one at a time, as well as incorporating nonlinearities. We observe that our predictions are considerably robust in most cases, allowing us to extend our conclusions to more practical learning setups.

海森矩阵损失曲面泛化性优化分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。