通过增强标注提升对抗训练中的类别鲁棒性差距
Narrowing Class-Wise Robustness Gaps in Adversarial Training
- 引入增强标注策略优化对抗训练过程
- 对抗鲁棒性提升53.50%,类别不平衡降低5.73%
- 适合关注模型公平性和泛化能力的研究者
为应对数据分布变化导致的准确率下降,研究常采用数据增强策略。对抗训练是其中一种方法,旨在提升对最坏情况分布偏移(由对抗样本引起)的鲁棒性。尽管该方法可增强鲁棒性,但可能损害对干净样本的泛化能力,并加剧不同类别间的性能差异。本文探讨了对抗训练对整体及类别特定性能的影响及其溢出效应。实验发现,训练期间增强标注可使对抗鲁棒性提升53.50%,缓解类别不平衡5.73%,从而在干净和对抗场景下均实现更高准确率,优于标准对抗训练。
原文摘要 · Abstract (English)
Efforts to address declining accuracy as a result of data shifts often involve various data-augmentation strategies. Adversarial training is one such method, designed to improve robustness to worst-case distribution shifts caused by adversarial examples. While this method can improve robustness, it may also hinder generalization to clean examples and exacerbate performance imbalances across different classes. This paper explores the impact of adversarial training on both overall and class-specific performance, as well as its spill-over effects. We observe that enhanced labeling during training boosts adversarial robustness by 53.50% and mitigates class imbalances by 5.73%, leading to improved accuracy in both clean and adversarial settings compared to standard adversarial training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。