arXiv:2512.14333cs.CVcs.AI2025-12被引 3

通过双重注意力机制干扰模型生成,有效防御文本误导性图像编辑。

Dual Attention Guided Defense Against Malicious Edits

  • 在多个时间步动态调节注意力分布,引导生成偏离目标区域。
  • 使注入噪声与模型预测噪声差异最大化,破坏生成一致性。
  • 对恶意编辑具有强鲁棒性,适合安全敏感的生成应用。

近期文本到图像扩散模型的发展使得基于文本提示的图像编辑成为可能,但也带来了伪造或有害内容滥用的伦理风险。现有防御方法通过嵌入不可察觉的扰动来缓解此问题,但对恶意篡改效果有限。为此,我们提出双注意力引导噪声扰动(DANP)免疫方法,通过添加不可察觉的扰动,破坏模型的语义理解与生成过程。DANP在多个时间步操作,同时影响跨注意力图与噪声预测过程,利用动态阈值生成掩码,识别与文本相关和无关区域,降低相关区域注意力、提升无关区域注意力,从而误导编辑至错误位置并保护目标。此外,该方法最大化注入噪声与模型预测噪声间的差异,进一步干扰生成。通过同时作用于注意力与噪声预测机制,DANP展现出优异的抗恶意编辑能力,大量实验验证其达到当前最优性能。

原文摘要 · Abstract (English)

Recent progress in text-to-image diffusion models has transformed image editing via text prompts, yet this also introduces significant ethical challenges from potential misuse in creating deceptive or harmful content. While current defenses seek to mitigate this risk by embedding imperceptible perturbations, their effectiveness is limited against malicious tampering. To address this issue, we propose a Dual Attention-Guided Noise Perturbation (DANP) immunization method that adds imperceptible perturbations to disrupt the model's semantic understanding and generation process. DANP functions over multiple timesteps to manipulate both cross-attention maps and the noise prediction process, using a dynamic threshold to generate masks that identify text-relevant and irrelevant regions. It then reduces attention in relevant areas while increasing it in irrelevant ones, thereby misguides the edit towards incorrect regions and preserves the intended targets. Additionally, our method maximizes the discrepancy between the injected noise and the model's predicted noise to further interfere with the generation. By targeting both attention and noise prediction mechanisms, DANP exhibits impressive immunity against malicious edits, and extensive experiments confirm that our method achieves state-of-the-art performance.

图像生成对抗防御扩散模型注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。