在潜在空间扰动图像,让模型无法学习用户数据,同时保持视觉清晰。
Latent Diffusion Unlearning: Protecting Against Unauthorized Personalization Through Trajectory Shifted Perturbations
- 在扩散模型的潜在空间中修改去噪轨迹起点,实现隐蔽扰动。
- 相比像素空间方法,感知质量提升8%-10%(PSNR/SSIM/FID)。
- 适合保护隐私数据,防止未经授权的个性化生成。
文本到图像的扩散模型在仅用少量用户图像的情况下,即可实现高效高保真的个性化。然而,这种能力也引发了数据隐私、知识产权保护和未经授权使用的问题。为缓解此类风险,利用图像投毒生成“不可学习”训练样本的方法应运而生。现有方法多在像素空间操作,导致图像出现噪声和伪影,难以察觉。本文提出一种新型基于模型的扰动策略,作用于扩散模型的潜在空间。该方法通过交替执行去噪与反演,并调整去噪轨迹的起始点,确保扰动后图像在视觉上与原图高度一致,同时对下游生成模型的反演和个性化具有强抵抗力。该方法将不可学习性融入潜在扩散模型框架,实现了可感知且有效的防御机制。我们在四个基准数据集上验证了该方法对前沿反演攻击的鲁棒性。结果表明,该方法在感知质量上显著提升(约8%-10%,涵盖PSNR、SSIM、FID),在五种对抗场景下平均鲁棒性提升约10%,有效保障敏感数据安全。
原文摘要 · Abstract (English)
Text-to-image diffusion models have demonstrated remarkable effectiveness in rapid and high-fidelity personalization, even when provided with only a few user images. However, the effectiveness of personalization techniques has lead to concerns regarding data privacy, intellectual property protection, and unauthorized usage. To mitigate such unauthorized usage and model replication, the idea of generating ``unlearnable'' training samples utilizing image poisoning techniques has emerged. Existing methods for this have limited imperceptibility as they operate in the pixel space which results in images with noise and artifacts. In this work, we propose a novel model-based perturbation strategy that operates within the latent space of diffusion models. Our method alternates between denoising and inversion while modifying the starting point of the denoising trajectory: of diffusion models. This trajectory-shifted sampling ensures that the perturbed images maintain high visual fidelity to the original inputs while being resistant to inversion and personalization by downstream generative models. This approach integrates unlearnability into the framework of Latent Diffusion Models (LDMs), enabling a practical and imperceptible defense against unauthorized model adaptation. We validate our approach on four benchmark datasets to demonstrate robustness against state-of-the-art inversion attacks. Results demonstrate that our method achieves significant improvements in imperceptibility ($\sim 8 \% -10\%$ on perceptual metrics including PSNR, SSIM, and FID) and robustness ( $\sim 10\%$ on average across five adversarial settings), highlighting its effectiveness in safeguarding sensitive data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。