arXiv:2502.16167cs.CVcs.AI2025-02被引 1

用后门机制防止文本生成图像模型被恶意个性化,保护隐私与版权。

PersGuard: Preventing Malicious Personalization in Text-to-Image Diffusion Models via Model Backdoors

  • 在模型发布前植入可触发的保护后门,确保盗用时生成特定输出。
  • 实验显示在灰盒、黑盒及多目标场景下,防护效果优于传统干扰方法。
  • 适合关注图像生成模型版权与隐私安全的研究者与开发者。

扩散模型(DMs)在文本到图像(T2I)生成方面取得进展,但其个性化能力引发严重隐私与版权问题。恶意用户可利用模型生成未经授权的人物肖像或艺术风格复制品。现有主动防御方法主要依赖对参考图像添加对抗性扰动以干扰训练,但存在局限:假设所有训练图像均预先扰动,且在数据集包含未扰动图像或经轻微变换时易失效。本文提出PersGuard,一种基于后门的新型框架,用于防止预训练T2I扩散模型被未经授权个性化。不同于扰动方法,我们假设保护者可在模型发布前嵌入保护后门。若下游用户使用受保护图像微调模型,模型将保留后门并生成预定义保护输出;而对于未受保护图像,后门在微调过程中被有效移除,以保障正常生成能力。我们将后门注入建模为统一优化问题,包含三项目标:后门行为损失(激活保护)、先验保持损失(维持标准生成能力)和新颖的后门保留损失。该保留损失专门设计以模拟个性化损失,确保后门在下游微调中仍具鲁棒性。大量实验在灰盒与黑盒设置、多对象保护及人脸身份保护场景下表明,PersGuard相比现有扰动方法提供更优的隐私保护效果。

原文摘要 · Abstract (English)

Diffusion models (DMs) have advanced text-to-image (T2I) synthesis, yet their personalization capabilities raise serious privacy and copyright concerns. Malicious actors can misuse these models to generate unauthorized portraits or artistic style replicas. Existing proactive defenses primarily rely on applying adversarial perturbations to reference images to disrupt training. However, these approaches face limitations: they assume all training images are pre-perturbed and are prone to failure when datasets contain unperturbed images or undergo minor data transformations. In this paper, we introduce PersGuard, a novel backdoor-based framework designed to prevent unauthorized personalization of pre-trained T2I diffusion models. Unlike perturbation-based methods, we assume protectors can embed protective backdoors into the models before their release. This mechanism ensures that if a downstream user fine-tunes the model on protected images, the model retains the backdoor and generates predefined protective outputs; conversely, for unprotected images, the backdoor is effectively removed during fine-tuning to ensure normal model utility. We formulate the backdoor injection as a unified optimization problem incorporating three objectives: a backdoor behavior loss to activate protection, a prior preservation loss to maintain standard generation capabilities, and a novel backdoor retention loss. The retention loss is specifically designed to mirror personalization loss, ensuring the backdoor remains robust during downstream fine-tuning. Extensive experiments across gray-box and black-box settings, multi-object protection, and facial identity protection demonstrate that PersGuard provides superior privacy protection compared to existing perturbation-based methods.

扩散模型隐私保护后门检测版权防护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。