通过协作植入隐蔽触发器,防止扩散模型被恶意编辑图像。
GuardDoor: Safeguarding Against Malicious Diffusion Editing via Protective Backdoors
- 模型方微调编码器嵌入保护后门,图像所有者可添加不可见触发器。
- 经压缩、加噪等预处理后仍能有效阻止非法编辑,输出无意义。
- 适合内容创作者与模型服务商合作保护数字资产的场景。
扩散模型的普及虽推动了图像编辑发展,但也引发未经授权修改带来的虚假信息和抄袭等问题。现有防御方法依赖对抗扰动,但易被压缩、加噪等简单预处理手段消除。为此,我们提出GuardDoor,一种新型且鲁棒的保护机制,促进图像所有者与模型提供商协作:模型方微调图像编码器以嵌入保护后门,图像所有者可为自有图像附加不可感知的触发器。当未经授权用户使用该扩散模型编辑受保护图像时,模型将生成无意义输出,从而降低恶意编辑风险。本方法在面对图像预处理操作时表现出更强鲁棒性,且具备大规模部署潜力。该工作凸显了模型提供商与图像所有者协同构建安全框架,在生成式AI时代保障数字内容安全的重要价值。
原文摘要 · Abstract (English)
The growing accessibility of diffusion models has revolutionized image editing but also raised significant concerns about unauthorized modifications, such as misinformation and plagiarism. Existing countermeasures largely rely on adversarial perturbations designed to disrupt diffusion model outputs. However, these approaches are found to be easily neutralized by simple image preprocessing techniques, such as compression and noise addition. To address this limitation, we propose GuardDoor, a novel and robust protection mechanism that fosters collaboration between image owners and model providers. Specifically, the model provider participating in the mechanism fine-tunes the image encoder to embed a protective backdoor, allowing image owners to request the attachment of imperceptible triggers to their images. When unauthorized users attempt to edit these protected images with this diffusion model, the model produces meaningless outputs, reducing the risk of malicious image editing. Our method demonstrates enhanced robustness against image preprocessing operations and is scalable for large-scale deployment. This work underscores the potential of cooperative frameworks between model providers and image owners to safeguard digital content in the era of generative AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。