arXiv:2506.17809cs.LGstat.ML2025-06被引 1

用软秩衡量损失曲面平坦度,更准确预测模型泛化性能。

Flatness After All?

  • 提出基于赫森矩阵软秩的平坦度度量方法
  • 在模型校准条件下,可精确捕捉泛化误差期望
  • 适用于非校准模型,比传统方法更稳健

深度学习泛化能力研究关注损失函数极小值处的曲率与泛化性能的关系,尤其在过参数化神经网络中。已有研究表明‘平坦’极小值通常比‘尖锐’极小值泛化更好。然而,也有实证显示深度网络即使在任意尖锐度下仍能良好泛化,这质疑了传统曲率度量的有效性。本文提出采用赫森矩阵的软秩作为平坦度衡量标准。当指数族神经网络模型完全校准时,且预测误差与其输出的一阶、二阶导数无相关性时,该度量能准确反映渐近期望泛化差距。对于非校准模型,该度量与著名的Takeuchi信息准则相关,仍可对不过度自信的模型提供可靠的泛化差距估计。实验表明,该方法相比基线更具鲁棒性。

原文摘要 · Abstract (English)

Recent literature generalization in deep learning has examined the relationship between the curvature of the loss function at minima and generalization, mainly in the context of overparameterized neural networks. A key observation is that "flat" minima tend to generalize better than "sharp" minima. While this idea is supported by empirical evidence, it has also been shown that deep networks can generalize even with arbitrary sharpness, as measured by either the trace or the spectral norm of the Hessian. In this paper, we argue that generalization could be assessed by measuring flatness using a soft rank measure of the Hessian. We show that when an exponential family neural network model is exactly calibrated, and its prediction error and its confidence on the prediction are not correlated with the first and the second derivative of the network's output, our measure accurately captures the asymptotic expected generalization gap. For non-calibrated models, we connect a soft rank based flatness measure to the well-known Takeuchi Information Criterion and show that it still provides reliable estimates of generalization gaps for models that are not overly confident. Experimental results indicate that our approach offers a robust estimate of the generalization gap compared to baselines.

泛化分析平坦度度量赫森矩阵模型校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。