揭露扩散净化防御的脆弱性,证明其易被梯度攻击突破
DiffBreak: Is Diffusion-Based Purification Robust?
- 通过理论分析发现扩散模型梯度可被攻击利用,导致净化输出仍具对抗性
- 实验显示单次净化评估会夸大防御效果,正确评估需考虑多次净化结果
- 提出新工具DiffBreak和多数投票机制,显著提升测试可靠性
基于扩散模型的净化(DBP)被视为抵御对抗样本的可靠方法,因其能将对抗样本投影到自然数据流形。本文推翻这一核心观点,理论证明梯度攻击可有效针对扩散模型而非分类器,导致净化输出与对抗分布对齐。这暴露出两个关键缺陷:错误的梯度计算和仅测试单次随机净化的评估协议。我们表明,在正确处理随机性和重提交风险后,DBP会失效。为此,我们提出DiffBreak——首个能可靠对DBP求导的工具包,消除了此前夸大鲁棒性的梯度缺陷。我们还分析了当前依赖单次净化的分类方案,指出其内在无效性,并提出基于统计多数投票(MV)的替代方案,实现部分但有意义的鲁棒性提升。进一步提出一种对抗深度伪造水印的优化方法,生成系统性扰动,即使在MV下也能击败DBP,挑战其可行性。
原文摘要 · Abstract (English)
Diffusion-based purification (DBP) has become a cornerstone defense against adversarial examples (AEs), regarded as robust due to its use of diffusion models (DMs) that project AEs onto the natural data manifold. We refute this core claim, theoretically proving that gradient-based attacks effectively target the DM rather than the classifier, causing DBP's outputs to align with adversarial distributions. This prompts a reassessment of DBP's robustness, attributing it to two critical flaws: incorrect gradients and inappropriate evaluation protocols that test only a single random purification of the AE. We show that with proper accounting for stochasticity and resubmission risk, DBP collapses. To support this, we introduce DiffBreak, the first reliable toolkit for differentiation through DBP, eliminating gradient flaws that previously further inflated robustness estimates. We also analyze the current defense scheme used for DBP where classification relies on a single purification, pinpointing its inherent invalidity. We provide a statistically grounded majority-vote (MV) alternative that aggregates predictions across multiple purified copies, showing partial but meaningful robustness gain. We then propose a novel adaptation of an optimization method against deepfake watermarking, crafting systemic perturbations that defeat DBP even under MV, challenging DBP's viability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。