arXiv:2602.01718cs.LG2026-02中稿 · TMLR被引 1

在图像退化场景下,通用性度量效果因环境而异,需按场景选择。

Generalization Measures under Controlled Covariate Shift: A Regime-Aware Benchmark

  • 在受控退化数据集CIFAR-10-C/P上测试多种泛化度量
  • 梯度和尖锐度相关度量在多数退化场景中表现最佳
  • 度量有效性依赖具体退化类型,不能通用

在深度学习中,预测模型泛化能力仍面临挑战。以往研究多聚焦于独立同分布(IID)场景,而本文在受控退化与扰动设置下重新评估图像分类器的泛化度量。使用CIFAR-10-C/P数据集,保持标签空间和任务不变,仅对输入图像施加退化或扰动。该设定使我们能重审Dziugaite等(2020)提出的鲁棒性质疑:泛化度量的有效性可能高度依赖实验条件。实验表明,不同退化环境下,泛化度量的效用差异显著。在三种CNN架构中,基于输入梯度和尖锐度的度量表现突出,部分家族度量结果接近且受架构影响。优化类、信息准则及尖锐度度量在相关性或局部可靠性分析中提供额外信号。结论指出,模型选择不应仅依赖传统IID环境下有效的度量,而应视具体退化场景,将泛化度量视为依赖特定环境的排序信号。

原文摘要 · Abstract (English)

Predicting generalization from quantities available before target-test evaluation remains a central challenge in deep learning. The systematic benchmark of Jiang et al. (2020) evaluated many generalization measures, but it focused on independent and identically distributed (IID) settings. We revisit this problem for image classifiers evaluated under controlled corruptions and perturbations. Our study uses CIFAR-10-C/P, where the label space and task remain fixed while the input images are degraded or perturbed. This setting also allows us to revisit the robustness concerns raised by Dziugaite et al. (2020), who showed that the apparent reliability of generalization measures can depend strongly on experimental conditions. Our experiments show that the usefulness of generalization measures is strongly regime-dependent. In our exploratory decision analysis across three CNN-style architectures, sharpness- and input-gradient-based measures are among the leading individual signals, whereas family results are close and architecture dependent. Optimization-based measures, Information Criteria, and Sharpness-based measures provide additional regime-dependent signals in correlation or local-reliability analyses. Together, these findings suggest that model selection should not rely only on measures favored by IID evaluation. Instead, within the evaluated CIFAR-10-C/P setting and architectures, generalization measures should be treated as regime-dependent ranking signals whose utility must be evaluated for the intended corruption or perturbation setting.

泛化能力模型评估退化测试基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。