用原型引导技术更可靠地消除图像模型中的宽泛概念。
Prototype-Guided Concept Erasure in Diffusion Models
- 通过聚类提取概念原型,作为生成时的负向条件
- 在多个基准上实现宽泛概念的可靠移除
- 适合需要可控安全生成的AI应用开发者
概念擦除广泛应用于图像生成中,以防止文本到图像模型生成不希望的内容。现有方法能有效消除具体、明确的概念(如皮卡丘或埃隆·马斯克),但在处理如“性”或“暴力”等宽泛概念时性能下降,因其范围广、多面性强,难以可靠擦除。为此,我们利用模型内在嵌入几何结构,识别编码特定概念的潜在嵌入,并通过聚类得到一组概念原型,总结模型对概念的内部表征,将其作为推理时的负向条件信号,实现精确可靠的擦除。大量实验表明,该方法在多个基准测试中显著提升了对宽泛概念的可靠移除能力,同时保持整体图像质量,为更安全、可控制的图像生成迈出关键一步。
原文摘要 · Abstract (English)
Concept erasure is extensively utilized in image generation to prevent text-to-image models from generating undesired content. Existing methods can effectively erase narrow concepts that are specific and concrete, such as distinct intellectual properties (e.g. Pikachu) or recognizable characters (e.g. Elon Musk). However, their performance degrades on broad concepts such as ``sexual'' or ``violent'', whose wide scope and multi-faceted nature make them difficult to erase reliably. To overcome this limitation, we exploit the model's intrinsic embedding geometry to identify latent embeddings that encode a given concept. By clustering these embeddings, we derive a set of concept prototypes that summarize the model's internal representations of the concept, and employ them as negative conditioning signals during inference to achieve precise and reliable erasure. Extensive experiments across multiple benchmarks show that our approach achieves substantially more reliable removal of broad concepts while preserving overall image quality, marking a step towards safer and more controllable image generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。