arXiv:2504.00038cs.LGcs.AI2025-04

用干净训练提升对抗训练效果,缓解泛化能力下降问题

Revisiting the Relationship between Adversarial and Clean Training: Why Clean Training Can Make Adversarial Training Better

  • 通过干净训练降低对抗训练的学习难度
  • 实验证明可提升先进对抗训练方法的性能
  • 适合关注模型鲁棒性与泛化平衡的研究者

对抗训练(AT)虽能有效提升模型的对抗鲁棒性,但常导致泛化能力下降。近期研究尝试用干净训练辅助对抗训练,但结论存在矛盾。本文系统总结代表性策略,基于多视角假设,统一解释不同研究间的矛盾现象。深入分析了以往研究中从干净训练模型传递到对抗训练模型的知识,发现可分为两类:降低学习难度和提供正确引导。基于此,提出新思路,利用干净训练进一步优化先进对抗训练方法。揭示了对抗训练泛化能力下降部分源于其难以学习某些样本特征,而充分结合干净训练可有效缓解该问题。

原文摘要 · Abstract (English)

Adversarial training (AT) is an effective technique for enhancing adversarial robustness, but it usually comes at the cost of a decline in generalization ability. Recent studies have attempted to use clean training to assist adversarial training, yet there are contradictions among the conclusions. We comprehensively summarize the representative strategies and, with a focus on the multi - view hypothesis, provide a unified explanation for the contradictory phenomena among different studies. In addition, we conduct an in - depth analysis of the knowledge combinations transferred from clean - trained models to adversarially - trained models in previous studies, and find that they can be divided into two categories: reducing the learning difficulty and providing correct guidance. Based on this finding, we propose a new idea of leveraging clean training to further improve the performance of advanced AT methods.We reveal that the problem of generalization degradation faced by AT partly stems from the difficulty of adversarial training in learning certain sample features, and this problem can be alleviated by making full use of clean training.

对抗训练泛化能力知识迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。