arXiv:2502.09352cs.LGcs.CV2025-02被引 5

提升深度神经网络对分布攻击的鲁棒性,不牺牲原有抗点攻击能力。

Wasserstein distributional adversarial training for deep neural networks

  • 基于Wasserstein距离构建分布鲁棒优化训练方法
  • 在RobustBench上验证可增强分布鲁棒性,保持点攻击鲁棒性
  • 适用于已有模型的高效微调,尤其对小数据集训练模型有效

针对深度神经网络的对抗攻击设计及其防御方法是当前研究热点。本文提出一种针对分布攻击威胁的训练方法,扩展了用于点攻击防御的TRADES方法。该方法基于近期进展,利用Wasserstein分布鲁棒优化问题的敏感性分析,提出一种高效的模型微调策略,可部署于已训练模型。我们在RobustBench上的多个预训练模型上测试该方法,结果表明额外训练显著提升了模型对Wasserstein分布攻击的鲁棒性,同时保持原有的点攻击鲁棒性,即使对于已非常成功的网络亦然。对于使用20-100M图像合成数据集预训练的模型,改进效果较弱;但令人惊讶的是,即便仅使用原始50k图像数据集训练的模型,本方法仍能带来性能提升。

原文摘要 · Abstract (English)

Design of adversarial attacks for deep neural networks, as well as methods of adversarial training against them, are subject of intense research. In this paper, we propose methods to train against distributional attack threats, extending the TRADES method used for pointwise attacks. Our approach leverages recent contributions and relies on sensitivity analysis for Wasserstein distributionally robust optimization problems. We introduce an efficient fine-tuning method which can be deployed on a previously trained model. We test our methods on a range of pre-trained models on RobustBench. These experimental results demonstrate the additional training enhances Wasserstein distributional robustness, while maintaining original levels of pointwise robustness, even for already very successful networks. The improvements are less marked for models pre-trained using huge synthetic datasets of 20-100M images. However, remarkably, sometimes our methods are still able to improve their performance even when trained using only the original training dataset (50k images).

对抗训练分布鲁棒Wasserstein微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。