arXiv:2504.14554cs.CRcs.CV2025-04被引 3

通过模型编辑实现精准文本到图像后门攻击,提升隐蔽性与成功率。

REDEditing: Relationship-Driven Precise Backdoor Poisoning on Text-to-Image Diffusion Models

论文配图:REDEditing: Relationship-Driven Precise Backdoor Poisoning on Text-to-Image Diffusion Models
图 1 · 摘自论文原文
  • 基于等价属性对齐与隐匿污染原则,实现概念重绑定的精准攻击。
  • 攻击成功率比现有方法高11%,仅用一行代码就提升自然度和隐蔽性24%。
  • 适合关注生成模型安全、后门防御的研究者与开发者参考。

生成式AI的快速发展凸显了文本到图像(T2I)模型安全的重要性,尤其是后门污染威胁。及时发现并缓解T2I模型中的安全漏洞,对确保生成模型的安全部署至关重要。本文探索了一种无需训练的新型后门污染范式——基于模型编辑技术,该技术近期被用于大语言模型的知识更新。然而,我们揭示了模型编辑技术对图像生成模型带来的潜在安全风险。在此工作中,我们建立了基于模型编辑的后门攻击原则,并提出关系驱动的精确后门污染方法REDEditing。基于等价属性对齐与隐匿污染原则,我们设计了等价关系检索与联合属性迁移方法,通过概念重绑定实现一致的后门图像生成。同时提出知识隔离约束以保障良性生成完整性。相比当前最优方法,本方法攻击成功率提升11%。值得注意的是,仅增加一行代码即可提升输出自然度,并使后门隐蔽性提高24%。本工作旨在提高对可编辑图像生成模型中此类安全漏洞的认识。

原文摘要 · Abstract (English)

The rapid advancement of generative AI highlights the importance of text-to-image (T2I) security, particularly with the threat of backdoor poisoning. Timely disclosure and mitigation of security vulnerabilities in T2I models are crucial for ensuring the safe deployment of generative models. We explore a novel training-free backdoor poisoning paradigm through model editing, which is recently employed for knowledge updating in large language models. Nevertheless, we reveal the potential security risks posed by model editing techniques to image generation models. In this work, we establish the principles for backdoor attacks based on model editing, and propose a relationship-driven precise backdoor poisoning method, REDEditing. Drawing on the principles of equivalent-attribute alignment and stealthy poisoning, we develop an equivalent relationship retrieval and joint-attribute transfer approach that ensures consistent backdoor image generation through concept rebinding. A knowledge isolation constraint is proposed to preserve benign generation integrity. Our method achieves an 11\% higher attack success rate compared to state-of-the-art approaches. Remarkably, adding just one line of code enhances output naturalness while improving backdoor stealthiness by 24\%. This work aims to heighten awareness regarding this security vulnerability in editable image generation models.

后门攻击图像生成模型编辑安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。