研究对抗扰动对分布鲁棒性的影响,发现适度偏倚数据上扰动能提升鲁棒性。
On the Effects of Adversarial Perturbations on Distribution Robustness
- 通过理论分析构建对抗训练的可计算替代方案
- 中等偏倚数据上ℓ∞扰动可提升分布鲁棒性
- 特征可分性决定鲁棒性增益是否成立,适合关注模型可靠性研究者
对抗鲁棒性指模型抵抗输入扰动的能力,而分布鲁棒性评估模型在数据分布漂移下的表现。尽管二者均旨在保证可靠性能,但已有研究表明两者存在权衡:对抗训练可能加剧对伪特征的依赖,损害分布鲁棒性,尤其在少数群体上的表现。本文通过分析在扰动数据上训练的模型,提供了一种可计算的对抗训练替代方案。除了已知的权衡外,研究还揭示了一个微妙现象:对中等偏倚数据施加ℓ∞扰动反而能提升分布鲁棒性。当数据高度偏斜时,若简单性偏差促使模型依赖核心特征(表现为更高特征可分性),该增益仍能维持。理论分析深化了对权衡机制的理解,强调特征可分性在其中的关键作用。尽管权衡在多数情况下依然存在,忽略特征可分性可能导致对鲁棒性的错误判断。
原文摘要 · Abstract (English)
Adversarial robustness refers to a model's ability to resist perturbation of inputs, while distribution robustness evaluates the performance of the model under data shifts. Although both aim to ensure reliable performance, prior work has revealed a tradeoff in distribution and adversarial robustness. Specifically, adversarial training might increase reliance on spurious features, which can harm distribution robustness, especially the performance on some underrepresented subgroups. We present a theoretical analysis of adversarial and distribution robustness that provides a tractable surrogate for per-step adversarial training by studying models trained on perturbed data. In addition to the tradeoff, our work further identified a nuanced phenomenon that $\ell_\infty$ perturbations on data with moderate bias can yield an increase in distribution robustness. Moreover, the gain in distribution robustness remains on highly skewed data when simplicity bias induces reliance on the core feature, characterized as greater feature separability. Our theoretical analysis extends the understanding of the tradeoff by highlighting the interplay of the tradeoff and the feature separability. Despite the tradeoff that persists in many cases, overlooking the role of feature separability may lead to misleading conclusions about robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。