arXiv:2410.12307cs.LGcs.CV2024-10NeurIPS被引 19

通过频域幅度混合提升模型对抗攻击鲁棒性

DAT: Improving Adversarial Robustness via Generative Amplitude Mix-up in Frequency Domain

  • 在频域中混合训练样本与干扰图像的幅度信息
  • 在多个数据集上显著增强对各类对抗攻击的防御能力
  • 适合关注模型安全性的研究人员和工程实践者

为保护深度神经网络免受对抗攻击,研究者提出了对抗训练(AT),将对抗样本融入模型训练。近期研究表明,对抗攻击对样本频谱相位中的模式影响更大,而这些模式通常包含关键语义信息,导致模型对对抗样本误分类。我们发现,若在对抗训练中将训练样本频谱的幅度与干扰图像的幅度进行混合,可引导模型关注不受对抗扰动影响的相位模式,从而提升鲁棒性。然而,如何选择合适的干扰图像仍具挑战——需在不改变相位的前提下混合幅度。为此,本文提出优化的对抗幅度生成器(AAG),实现鲁棒性提升与相位保持之间的更好平衡。基于该生成器及高效的对抗样本生成流程,设计了新的双对抗训练(DAT)策略。在多个数据集上的实验表明,所提DAT方法显著提升了模型对多样化对抗攻击的鲁棒性。

原文摘要 · Abstract (English)

To protect deep neural networks (DNNs) from adversarial attacks, adversarial training (AT) is developed by incorporating adversarial examples (AEs) into model training. Recent studies show that adversarial attacks disproportionately impact the patterns within the phase of the sample's frequency spectrum -- typically containing crucial semantic information -- more than those in the amplitude, resulting in the model's erroneous categorization of AEs. We find that, by mixing the amplitude of training samples' frequency spectrum with those of distractor images for AT, the model can be guided to focus on phase patterns unaffected by adversarial perturbations. As a result, the model's robustness can be improved. Unfortunately, it is still challenging to select appropriate distractor images, which should mix the amplitude without affecting the phase patterns. To this end, in this paper, we propose an optimized Adversarial Amplitude Generator (AAG) to achieve a better tradeoff between improving the model's robustness and retaining phase patterns. Based on this generator, together with an efficient AE production procedure, we design a new Dual Adversarial Training (DAT) strategy. Experiments on various datasets show that our proposed DAT leads to significantly improved robustness against diverse adversarial attacks.

对抗训练频域分析鲁棒性提升

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。