发现删除的概念仍可被生成,揭示扩散模型概念擦除存在漏洞。
Memories of Forgotten Concepts
- 用反演方法找到能生成被删除概念的潜在编码
- 被删概念图像可由大量不同潜在种子生成,且分布重叠外部图像
- 表明完全清除特定概念信息在技术上几乎不可能
扩散模型在文本到图像生成中占主导地位,但可能生成不当内容或私人数据。为缓解此问题,已有研究探索概念擦除技术以限制特定概念的生成。本文揭示,被擦除的概念信息仍存在于模型中,且通过合适的潜在表示可生成该概念的高质量图像。利用反演方法,我们发现存在能够生成被擦除概念图像的潜在种子。进一步表明,这些潜在编码的分布与非被擦除概念图像的分布存在重叠。我们还证明,对于每个被擦除概念中的图像,均可生成多个生成该概念的潜在种子。由于存在大量能生成被擦除概念图像的潜在空间,我们的结果表明完全擦除概念信息在理论上不可行,凸显了当前概念擦除技术的潜在安全风险。
原文摘要 · Abstract (English)
Diffusion models dominate the space of text-to-image generation, yet they may produce undesirable outputs, including explicit content or private data. To mitigate this, concept ablation techniques have been explored to limit the generation of certain concepts. In this paper, we reveal that the erased concept information persists in the model and that erased concept images can be generated using the right latent. Utilizing inversion methods, we show that there exist latent seeds capable of generating high quality images of erased concepts. Moreover, we show that these latents have likelihoods that overlap with those of images outside the erased concept. We extend this to demonstrate that for every image from the erased concept set, we can generate many seeds that generate the erased concept. Given the vast space of latents capable of generating ablated concept images, our results suggest that fully erasing concept information may be intractable, highlighting possible vulnerabilities in current concept ablation techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。