arXiv:2505.22839cs.LGcs.AI2025-05被引 1

揭秘扩散模型如何提升对抗鲁棒性:随机性与空间压缩是关键

Demystifying Adversarial Robustness in Diffusion Models: Compression, Randomness, and Geometry

  • 通过分解随机性和图像空间压缩,揭示扩散模型的抗扰动机制
  • 随机性导致梯度遮蔽,使攻击方向相似度下降,提升鲁棒性
  • 模型压缩能力与真实鲁棒性增益存在可解释的规律关系

近期研究指出,扩散模型显著提升了深度神经网络的对抗鲁棒性。尽管已有直观解释,但其内在机制仍不明确。本文发现,扩散模型反而增大了净化后图像与原始样本的ℓ_p距离,推翻了‘净化使扰动图像更接近干净样本’的假设。我们提出统一框架,将鲁棒性提升归因于两点:(i) 内部随机性引起的梯度遮蔽;(ii) 图像空间的压缩。实验表明,随机性导致的梯度遮蔽无法通过期望-变换(EOT)缓解,其效果由最优攻击方向与实际攻击方向的余弦相似度决定,符合超球冠模型预测。此外,在固定随机性下,扩散模型显著压缩图像空间,且压缩能力与真实鲁棒性增益呈可解释的正相关。理论分析进一步表明,收敛的得分场(score fields)解释了这一压缩效应。本研究揭示了扩散净化的新机制,为构建更有效、更合理的对抗净化系统提供指导。

原文摘要 · Abstract (English)

Recent studies suggest that diffusion models significantly improve the empirical adversarial robustness of deep neural network models. While intuitive explanations have been proposed, the mechanisms underlying diffusion-based robustness remain largely unclear. This work aims to demystify how diffusion models improve adversarial robustness. We observe that diffusion models surprisingly increase the $\ell_p$ distance to clean samples, thus rejecting the hypothesis that purification denoises perturbed images closer to the clean ones. Next, we provide a unifying account of the robustness improvement in diffusion-based purification by decomposing it into two sources: (i) gradient masking induced by randomness; (ii) compression of the image space. First, we find that the purified images are heavily influenced by the internal randomness of diffusion models. This randomness leads to gradient masking that cannot be removed by the previously proposed remedy, i.e., expectation-over-transformation (EOT). The improvement in robustness due to randomness is determined by the cosine similarity of the optimal vs. empirical attack directions, as predicted by a hyperspherical cap model of the adversarial regions. Second, we find that, when fixing the randomness, diffusion models substantially compress the image space. Importantly, we discover a lawful relationship between the model's ability to compress the image space and the genuine adversarial robustness gain. Further theoretical analyses show that convergent score fields encoded in diffusion models explain these compression effects. Our findings reveal new insights into the mechanisms underlying diffusion-based purification, and offer guidance for developing more effective and principled adversarial purification systems.

扩散模型对抗鲁棒性随机性空间压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。