用隐性水印触发图像编辑模型后门,攻击隐蔽且高效
Invisible Backdoor Triggers in Image Editing Model via Deep Watermarking
- 通过中毒训练数据嵌入不可见水印作为后门触发器
- 在不同水印模型下攻击成功率超90%,不影响正常编辑
- 适合研究模型安全或对抗攻击的开发者参考
扩散模型在图像生成与编辑方面取得显著进展,但近期研究揭示其易受后门攻击:输入中嵌入特定模式可操控模型行为。现有工作多聚焦图像生成环节,而针对图像编辑的后门攻击仍不充分。少数研究虽涉及编辑场景,但普遍采用可见触发器,导致输入图像出现明显异常,实用性差。本文提出一种新攻击框架,通过中毒训练数据将不可见触发器嵌入图像编辑过程。利用现成深度水印模型,将人眼无法察觉的水印编码为后门触发器。目标是使模型在接收到带水印输入时生成预设后门目标,而对正常图像仍按提示进行常规编辑。在多种水印模型上的大量实验表明,该方法实现高攻击成功率。此外,关于水印特征与后门攻击效果的分析进一步验证了方法的有效性。代码已开源。
原文摘要 · Abstract (English)
Diffusion models have achieved remarkable progress in both image generation and editing. However, recent studies have revealed their vulnerability to backdoor attacks, in which specific patterns embedded in the input can manipulate the model's behavior. Most existing research in this area has proposed attack frameworks focused on the image generation pipeline, leaving backdoor attacks in image editing relatively unexplored. Among the few studies targeting image editing, most utilize visible triggers, which are impractical because they introduce noticeable alterations to the input image before editing. In this paper, we propose a novel attack framework that embeds invisible triggers into the image editing process via poisoned training data. We leverage off-the-shelf deep watermarking models to encode imperceptible watermarks as backdoor triggers. Our goal is to make the model produce the predefined backdoor target when it receives watermarked inputs, while editing clean images normally according to the given prompt. With extensive experiments across different watermarking models, the proposed method achieves promising attack success rates. In addition, the analysis results of the watermark characteristics in term of backdoor attack further support the effectiveness of our approach. The code is available at:https://github.com/aiiu-lab/BackdoorImageEditing
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。