arXiv:2511.19910eess.IVcs.CV2025-11被引 1

提出双层防御框架,同时对抗微调和零样本定制扩散模型的隐私泄露风险。

DLADiff: A Dual-Layer Defense Framework against Fine-Tuning and Zero-Shot Customization of Diffusion Models

  • 采用双代理模型与动态微调结合,抵御未经授权的模型微调。
  • 在仅用3~5张图微调时,防御成功率超90%;零样本生成防御效果显著提升。
  • 适合关注人脸隐私保护的AI安全研究者及应用开发者使用。

随着扩散模型的快速发展,多种微调方法被提出,仅需3至5张训练图像即可生成高度逼真的目标内容图像。近期,无需修改模型权重的零样本生成方法也出现,仅凭一张参考图像即可生成高度真实的输出。然而,这些技术也带来了严重的面部隐私风险:攻击者可利用少量甚至单张人脸图像,通过模型定制生成几乎与原身份一致的合成身份。尽管已有研究关注扩散模型定制的防御,但多数方法仅针对微调,忽视零样本生成的防护。为此,本文提出双层反扩散防御框架DLADiff,兼顾两类攻击。第一层通过双代理模型(DSUR)与交替动态微调(ADFT)机制,融合对抗训练与预微调模型先验知识,有效防御未经授权的微调。第二层设计简洁却高效,显著抑制零样本生成。大量实验表明,该方法在抵御微调方面显著优于现有方法,并在零样本生成防护上达到前所未有的效果。

原文摘要 · Abstract (English)

With the rapid advancement of diffusion models, a variety of fine-tuning methods have been developed, enabling high-fidelity image generation with high similarity to the target content using only 3 to 5 training images. More recently, zero-shot generation methods have emerged, capable of producing highly realistic outputs from a single reference image without altering model weights. However, technological advancements have also introduced significant risks to facial privacy. Malicious actors can exploit diffusion model customization with just a few or even one image of a person to create synthetic identities nearly identical to the original identity. Although research has begun to focus on defending against diffusion model customization, most existing defense methods target fine-tuning approaches and neglect zero-shot generation defenses. To address this issue, this paper proposes Dual-Layer Anti-Diffusion (DLADiff) to defense both fine-tuning methods and zero-shot methods. DLADiff contains a dual-layer protective mechanism. The first layer provides effective protection against unauthorized fine-tuning by leveraging the proposed Dual-Surrogate Models (DSUR) mechanism and Alternating Dynamic Fine-Tuning (ADFT), which integrates adversarial training with the prior knowledge derived from pre-fine-tuned models. The second layer, though simple in design, demonstrates strong effectiveness in preventing image generation through zero-shot methods. Extensive experimental results demonstrate that our method significantly outperforms existing approaches in defending against fine-tuning of diffusion models and achieves unprecedented performance in protecting against zero-shot generation.

隐私保护扩散模型防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。