提出空间注意力遗忘技术,精准清除扩散模型中的后门攻击
Backdoor Defense in Diffusion Models via Spatial Attention Unlearning
- 通过潜空间操作与空间注意力机制定位并移除后门触发信号
- 对多种触发类型实现100%清除率,生成图像CLIP分数达0.7023
- 适合关注图像生成安全性的研究人员和实际部署开发者
文本到图像扩散模型正面临后门攻击威胁,恶意训练数据会导致模型在特定触发器出现时生成异常输出。尽管分类模型已有丰富防御手段,生成模型因高维输出空间难以检测细微扰动,其防御仍不充分。本文提出空间注意力遗忘(SAU)技术,利用潜空间操控与空间注意力机制,精准隔离并移除后门触发的潜在表示,实现高效去毒。我们在多种后门攻击(包括像素级与风格级触发)上验证了SAU的有效性,实现100%触发消除率。同时,生成图像的CLIP分数达0.7023,优于现有方法,且保持高质量、语义一致的生成能力。结果表明,SAU是抵御扩散模型后门攻击的鲁棒、可扩展且实用的解决方案。
原文摘要 · Abstract (English)
Text-to-image diffusion models are increasingly vulnerable to backdoor attacks, where malicious modifications to the training data cause the model to generate unintended outputs when specific triggers are present. While classification models have seen extensive development of defense mechanisms, generative models remain largely unprotected due to their high-dimensional output space, which complicates the detection and mitigation of subtle perturbations. Defense strategies for diffusion models, in particular, remain under-explored. In this work, we propose Spatial Attention Unlearning (SAU), a novel technique for mitigating backdoor attacks in diffusion models. SAU leverages latent space manipulation and spatial attention mechanisms to isolate and remove the latent representation of backdoor triggers, ensuring precise and efficient removal of malicious effects. We evaluate SAU across various types of backdoor attacks, including pixel-based and style-based triggers, and demonstrate its effectiveness in achieving 100% trigger removal accuracy. Furthermore, SAU achieves a CLIP score of 0.7023, outperforming existing methods while preserving the model's ability to generate high-quality, semantically aligned images. Our results show that SAU is a robust, scalable, and practical solution for securing text-to-image diffusion models against backdoor attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。