arXiv:2505.24592cs.LGcs.AI2025-05AAAI被引 1

从平坦极小值视角揭示数据增强如何提升模型鲁棒性

A Flat Minima Perspective on Understanding Augmentations and Model Robustness

  • 基于平坦极小值理论,解析标签保持型增强的鲁棒性机制
  • 在CIFAR和ImageNet上验证增强对多种分布偏移的鲁棒性提升
  • 首次统一解释各类增强对不同分布偏移的增益效果

模型鲁棒性指模型在未知分布偏移下仍能良好泛化的能力,包括数据损坏和对抗攻击。数据增强是提升鲁棒性的主流且有效手段。尽管各类增强在不同领域表现优异,但对其提升鲁棒性的统一理论理解仍缺失。本文从平坦极小值与泛化界视角,理论上揭示了标签保持型增强在应对多样分布偏移时带来鲁棒性的普遍条件,该条件在实践中与多种分布偏移下的鲁棒性高度相关。不同于多数早期工作,本理论框架适用于所有标签保持型增强,不限于特定分布偏移。我们在CIFAR和ImageNet基准上的常见损坏与对抗鲁棒性测试中验证了理论有效性。

原文摘要 · Abstract (English)

Model robustness indicates a model's capability to generalize well on unforeseen distributional shifts, including data corruptions and adversarial attacks. Data augmentation is one of the most prevalent and effective ways to enhance robustness. Despite the great success of the diverse augmentations in different fields, a unified theoretical understanding of their efficacy in improving model robustness is lacking. We theoretically reveal a general condition for label-preserving augmentations to bring robustness to diverse distribution shifts through the lens of flat minima and generalization bound, which de facto turns out to be strongly correlated with robustness against different distribution shifts in practice. Unlike most earlier works, our theoretical framework accommodates all the label-preserving augmentations and is not limited to particular distribution shifts. We substantiate our theories through different simulations on the existing common corruption and adversarial robustness benchmarks based on the CIFAR and ImageNet datasets.

数据增强鲁棒性平坦极小值

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。