arXiv:2510.07758cs.LG2025-10被引 1

用瑞尼熵定义新尖锐度,更好预测模型泛化能力

Rényi Sharpness: A Novel Sharpness that Strongly Correlates with Generalization

  • 基于损失海森矩阵谱的不均匀性,提出瑞尼尖锐度
  • 在多种场景下与泛化性能强相关,优于传统指标
  • 可作正则项提升训练效果,适合模型优化研究者

尖锐度(损失极小值处的海森矩阵性质)被认为是神经网络泛化能力的良好指标。然而,现有尖锐度度量与泛化之间的相关性不够强,有时甚至矛盾。本文关键观察是:泛化真正关键的是损失海森矩阵谱的平均不均匀性。传统度量如迹尖锐度(tr(H))关注谱的平均值,最大特征值尖锐度(λ_max(H))关注最大波动,均不足以准确预测泛化。为此,本文引入信息论中的瑞尼熵概念,将其扩展至非负向量(如海森谱),定义新的瑞尼尖锐度为损失海森矩阵谱瑞尼熵的负值。大量实验表明,瑞尼尖锐度在不同场景下与泛化性能呈现强而一致的相关性。进一步,利用其优良的重参数不变性,建立了两个关于瑞尼尖锐度的泛化界。最后,首次尝试将瑞尼尖锐度用于正则化,提出瑞尼尖锐度感知最小化(RSAM),使用其变体作为正则项。实验显示,RSAM在性能上媲美当前最优的SAM算法,显著优于基于最大特征值的常规SAM。

原文摘要 · Abstract (English)

Sharpness (of the loss minima) is widely believed to be a good indicator of generalization of neural networks. Unfortunately, the correlation between existing sharpness measures and generalization is not as strong as expected, and sometimes even contradiction occurs. To address this problem, a key observation in this paper is: what really matters for generalization is the average spread (or unevenness) of the spectrum of loss Hessian $\mathbf{H}$. For this reason, conventional sharpness measures, such as trace sharpness $\operatorname{tr}(\mathbf{H})$, which cares about the average value of the spectrum, or max-eigenvalue sharpness $λ_{\max}(\mathbf{H})$, which concerns the maximum spread of the spectrum, are not sufficient to well predict generalization. To characterize the average spread of the Hessian spectrum, we leverage the notion of Rényi entropy in information theory, which captures the unevenness of a probability vector and can thus be extended to a general non-negative vector, such as the Hessian spectrum at loss minima. Specifically, we propose Rényi sharpness, defined as the negative of the Rényi entropy of loss Hessian $\mathbf{H}$. Extensive experiments demonstrate that Rényi sharpness exhibits strong and consistent correlation with generalization in various scenarios. Moreover, two generalization bounds with respect to Rényi sharpness are established by exploiting its desirable reparametrization invariance property. Finally, as an initial attempt to exploit Rényi sharpness for regularization, Rényi Sharpness Aware Minimization (RSAM) is proposed, where a variant of Rényi sharpness is used as the regularizer. RSAM is competitive with state-of-the-art SAM algorithms and far better than conventional SAM based on max-eigenvalue sharpness.

尖锐度泛化分析优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。