arXiv:2503.16835cs.CVcs.LG2025-03被引 6

通过子空间投影彻底移除扩散模型中的不想要概念

Safe and Reliable Diffusion Models via Subspace Projection

  • 用文本嵌入的低维结构定位目标概念子空间
  • 投影消除提示词中概念信息,防止变相重现
  • 适合需要内容安全可控的图像生成场景

大规模文生图扩散模型虽能生成高质量图像,但可能无意生成版权作品或不当内容。现有方法常无法完全消除特定概念,导致其以微妙形式重现——例如屏蔽‘梵高’后仍可能生成《星夜》。本文提出SAFER方法,基于文本嵌入空间的低维结构,首先通过文本反演从参考图像学习目标概念的优化嵌入,进而识别其专属子空间Sc;再将提示词嵌入投影至Sc的补空间,实现概念擦除。此外引入子空间扩展策略确保全面清除。大量实验表明,SAFER可有效且一致地移除目标概念,同时保持生成质量。

原文摘要 · Abstract (English)

Large-scale text-to-image (T2I) diffusion models have revolutionized image generation, enabling the synthesis of highly detailed visuals from textual descriptions. However, these models may inadvertently generate inappropriate content, such as copyrighted works or offensive images. While existing methods attempt to eliminate specific unwanted concepts, they often fail to ensure complete removal, allowing the concept to reappear in subtle forms. For instance, a model may successfully avoid generating images in Van Gogh's style when explicitly prompted with 'Van Gogh', yet still reproduce his signature artwork when given the prompt 'Starry Night'. In this paper, we propose SAFER, a novel and efficient approach for thoroughly removing target concepts from diffusion models. At a high level, SAFER is inspired by the observed low-dimensional structure of the text embedding space. The method first identifies a concept-specific subspace $S_c$ associated with the target concept c. It then projects the prompt embeddings onto the complementary subspace of $S_c$, effectively erasing the concept from the generated images. Since concepts can be abstract and difficult to fully capture using natural language alone, we employ textual inversion to learn an optimized embedding of the target concept from a reference image. This enables more precise subspace estimation and enhances removal performance. Furthermore, we introduce a subspace expansion strategy to ensure comprehensive and robust concept erasure. Extensive experiments demonstrate that SAFER consistently and effectively erases unwanted concepts from diffusion models while preserving generation quality.

扩散模型内容安全子空间投影文本反演

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。