arXiv:2503.13945cs.CV2025-03

提出双阶段对抗攻击方法,提升扩散模型定制的防御效果。

Make the Most of Everything: Further Considerations on Disrupting Diffusion-based Customization

  • 分两阶段攻击:先生成对抗提示向量,再干扰图像生成过程
  • 在主流人脸数据集上实现10%-30%的抗定制性能提升
  • 适合研究隐私保护与生成模型安全性的研究人员

文本到图像扩散模型的微调技术虽可实现图像定制,但存在隐私泄露和观点操纵风险。现有研究多聚焦于提示或图像级对抗攻击,却忽视了二者关联及模型内部模块与输入的关系,导致实际威胁场景下防御效果受限。本文提出双阶段对抗攻击方法DADiff,首次将提示级对抗攻击融入图像级对抗样本生成过程。第一阶段生成提示级对抗向量以引导后续攻击;第二阶段在端到端攻击UNet模型的同时,干扰其自注意力与交叉注意力模块,破坏图像像素与提示间的关联性,并使实例提示与对抗提示向量在图像中产生的交叉注意力结果趋于一致。此外,引入局部随机时间步梯度集成策略,通过整合多个分割时间步的随机梯度更新对抗扰动。在多个主流人脸数据集上的实验表明,相比现有方法,DADiff在跨提示、关键词错位、跨模型和跨机制场景下的抗定制性能提升10%-30%。

原文摘要 · Abstract (English)

The fine-tuning technique for text-to-image diffusion models facilitates image customization but risks privacy breaches and opinion manipulation. Current research focuses on prompt- or image-level adversarial attacks for anti-customization, yet it overlooks the correlation between these two levels and the relationship between internal modules and inputs. This hinders anti-customization performance in practical threat scenarios. We propose Dual Anti-Diffusion (DADiff), a two-stage adversarial attack targeting diffusion customization, which, for the first time, integrates the adversarial prompt-level attack into the generation process of image-level adversarial examples. In stage 1, we generate prompt-level adversarial vectors to guide the subsequent image-level attack. In stage 2, besides conducting the end-to-end attack on the UNet model, we disrupt its self- and cross-attention modules, aiming to break the correlations between image pixels and align the cross-attention results computed using instance prompts and adversarial prompt vectors within the images. Furthermore, we introduce a local random timestep gradient ensemble strategy, which updates adversarial perturbations by integrating random gradients from multiple segmented timesets. Experimental results on various mainstream facial datasets demonstrate 10%-30% improvements in cross-prompt, keyword mismatch, cross-model, and cross-mechanism anti-customization with DADiff compared to existing methods.

对抗攻击扩散模型隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。