arXiv:2505.19425cs.CVcs.CR2025-05

通过扰动注意力查询,破坏扩散模型的结构生成能力,防止恶意修复。

Structure Disruption: Subverting Malicious Diffusion-Based Inpainting via Self-Attention Query Perturbation

  • 在去噪初始阶段扰动自注意力查询,破坏轮廓生成机制。
  • 在公开数据集上实现最优防护效果,保持强鲁棒性。
  • 适合保护敏感图像区域,防御社交媒体上的恶意编辑。

扩散模型在图像修复与编辑方面能力迅速提升,但也带来了严重的社会风险。攻击者可利用社交媒体中的用户图像生成误导或有害内容。尽管对抗扰动能干扰修复过程,但全局扰动方法在掩码引导的编辑任务中因空间限制而失效。为此,我们提出结构破坏攻击(SDA),一种强大的保护框架,用于防御针对敏感图像区域的基于扩散的编辑。基于扩散模型自注意力机制对轮廓的聚焦特性,SDA 在初始去噪步骤中优化扰动,破坏自注意力中的查询,从而摧毁轮廓生成过程。这种针对性干扰直接破坏了扩散模型的结构生成能力,有效阻止其生成连贯图像。我们通过可视化和在公开数据集上的大量实验验证了该方法的有效性,结果显示 SDA 在保持强鲁棒性的同时达到了当前最优的防护性能。

原文摘要 · Abstract (English)

The rapid advancement of diffusion models has enhanced their image inpainting and editing capabilities but also introduced significant societal risks. Adversaries can exploit user images from social media to generate misleading or harmful content. While adversarial perturbations can disrupt inpainting, global perturbation-based methods fail in mask-guided editing tasks due to spatial constraints. To address these challenges, we propose Structure Disruption Attack (SDA), a powerful protection framework for safeguarding sensitive image regions against inpainting-based editing. Building upon the contour-focused nature of self-attention mechanisms of diffusion models, SDA optimizes perturbations by disrupting queries in self-attention during the initial denoising step to destroy the contour generation process. This targeted interference directly disrupts the structural generation capability of diffusion models, effectively preventing them from producing coherent images. We validate our motivation through visualization techniques and extensive experiments on public datasets, demonstrating that SDA achieves state-of-the-art (SOTA) protection performance while maintaining strong robustness.

扩散模型图像安全对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。