arXiv:2504.11034cs.CV2025-04CVPR被引 1

用扩散模型净化图像,防御高低频对抗攻击

Defending Against Frequency-Based Attacks with Diffusion Models

  • 用扩散模型作为净化器,先去噪再分类
  • 在低频到高频区域均有效抑制对抗扰动
  • 适合应对未知攻击类型,泛化能力强

对抗训练是提升模型鲁棒性的常用策略,但通常仅针对特定攻击类型,难以泛化到未见威胁模型。对抗净化通过生成模型在分类前去除扰动,因其与分类器及威胁模型独立训练,更擅长处理未见过的攻击场景。扩散模型在噪声净化方面表现优异,不仅能应对像素级对抗扰动,还能缓解非对抗性数据偏移。本文将研究范围从像素级鲁棒性扩展至谱域和空间域的对抗攻击,发现净化方法在低频至高频区域均能有效缓解多种失真模式。

原文摘要 · Abstract (English)

Adversarial training is a common strategy for enhancing model robustness against adversarial attacks. However, it is typically tailored to the specific attack types it is trained on, limiting its ability to generalize to unseen threat models. Adversarial purification offers an alternative by leveraging a generative model to remove perturbations before classification. Since the purifier is trained independently of both the classifier and the threat models, it is better equipped to handle previously unseen attack scenarios. Diffusion models have proven highly effective for noise purification, not only in countering pixel-wise adversarial perturbations but also in addressing non-adversarial data shifts. In this study, we broaden the focus beyond pixel-wise robustness to explore the extent to which purification can mitigate both spectral and spatial adversarial attacks. Our findings highlight its effectiveness in handling diverse distortion patterns across low- to high-frequency regions.

对抗攻击扩散模型图像净化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。