arXiv:2505.20934cs.LG2025-05被引 4

用扩散模型生成更自然、迁移性更强的对抗样本。

NatADiff: Adversarial Boundary Guidance for Natural Adversarial Diffusion

  • 通过引导扩散过程逼近真实与对抗类的交集,生成结构合理的对抗样本。
  • 攻击成功率媲美顶尖方法,跨模型迁移率显著提升。
  • 生成样本更贴近真实测试误差,适合研究模型鲁棒性与防御机制。

对抗样本利用深度学习模型所学流形中的不规则性导致误分类。研究这些对抗样本可揭示模型分类所依赖的特征,从而提升对后续攻击的鲁棒性。然而,现有工作多聚焦于受限对抗样本,难以反映真实场景下的测试误差。为此,我们提出NatADiff,一种基于去噪扩散模型生成自然对抗样本的方法。该方法基于观察:自然对抗样本常包含对抗类的结构特征,模型可借此捷径分类而非真正区分类别。为此,我们引导扩散轨迹向真实与对抗类的交集靠近,结合时间回溯采样与增强分类器引导,在保持图像保真度的同时提升攻击迁移性。实验表明,该方法在攻击成功率上媲美当前最优技术,且在跨模型迁移性和与自然测试误差的契合度(以FID衡量)方面表现更优。结果表明,NatADiff生成的对抗样本不仅迁移能力更强,也更真实地模拟了实际测试中的错误模式。

原文摘要 · Abstract (English)

Adversarial samples exploit irregularities in the manifold `learned' by deep learning models to cause misclassifications. The study of these adversarial samples provides insight into the features a model uses to classify inputs, which can be leveraged to improve robustness against future attacks. However, much of the existing literature focuses on constrained adversarial samples, which do not accurately reflect test-time errors encountered in real-world settings. To address this, we propose `NatADiff', an adversarial sampling scheme that leverages denoising diffusion to generate natural adversarial samples. Our approach is based on the observation that natural adversarial samples frequently contain structural elements from the adversarial class. Deep learning models can exploit these structural elements to shortcut the classification process, rather than learning to genuinely distinguish between classes. To leverage this behavior, we guide the diffusion trajectory towards the intersection of the true and adversarial classes, combining time-travel sampling with augmented classifier guidance to enhance attack transferability while preserving image fidelity. Our method achieves comparable attack success rates to current state-of-the-art techniques, while exhibiting significantly higher transferability across model architectures and better alignment with natural test-time errors as measured by FID. These results demonstrate that NatADiff produces adversarial samples that not only transfer more effectively across models, but more faithfully resemble naturally occurring test-time errors.

对抗样本扩散模型迁移性鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。