arXiv:2603.09016cs.LGcs.CV2026-03

提出一种精确的卷积网络平坦度度量,可更好预测模型泛化性能。

An accurate flatness measure to estimate the generalization performance of CNN models

  • 基于交叉熵损失的海森迹闭式解,适配全局平均池化结构
  • 考虑卷积层参数缩放对称性与滤波器交互,具架构感知能力
  • 在图像分类任务中验证其对泛化性能评估的有效性

基于谱或海森迹的平坦度度量广泛用作深度网络泛化能力的代理指标。然而,现有方法多针对全连接结构,依赖海森迹的随机估计,或忽略现代卷积神经网络(CNN)特有的几何结构。本文针对使用全局平均池化后接线性分类器的广泛且实用的CNN类别,推导出交叉熵损失关于卷积核的海森迹闭式表达式。在此基础上,将相对平坦度概念特化至卷积层,得到一种参数化感知的平坦度度量,能正确建模卷积与池化带来的缩放对称性和滤波器交互。最后,我们在标准图像分类基准上对多种CNN进行了实证研究。结果表明,该度量可作为评估和比较CNN模型泛化性能的稳健工具,并指导实际中的架构与训练策略设计。

原文摘要 · Abstract (English)

Flatness measures based on the spectrum or the trace of the Hessian of the loss are widely used as proxies for the generalization ability of deep networks. However, most existing definitions are either tailored to fully connected architectures, relying on stochastic estimators of the Hessian trace, or ignore the specific geometric structure of modern Convolutional Neural Networks (CNNs). In this work, we develop a flatness measure that is both exact and architecturally faithful for a broad and practically relevant class of CNNs. We first derive a closed-form expression for the trace of the Hessian of the cross-entropy loss with respect to convolutional kernels in networks that use global average pooling followed by a linear classifier. Building on this result, we then specialize the notion of relative flatness to convolutional layers and obtain a parameterization-aware flatness measure that properly accounts for the scaling symmetries and filter interactions induced by convolution and pooling. Finally, we empirically investigate the proposed measure on families of CNNs trained on standard image-classification benchmarks. The results obtained suggest that the proposed measure can serve as a robust tool to assess and compare the generalization performance of CNN models, and to guide the design of architecture and training choices in practice.

卷积网络平坦度度量泛化性能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。