arXiv:2510.22851cs.CVcs.AI2025-10NeurIPS被引 8

无需训练即可精准删除文本中的有害概念,保持图像质量

Semantic Surgery: Zero-Shot Concept Erasure in Diffusion Models

  • 直接在文本嵌入阶段动态抵消目标概念影响
  • 物体删除完整率达93.58,显性内容降至1例,风格保留无损
  • 适合需要安全可控生成的AI应用开发者

文本到图像扩散模型中的概念擦除对缓解有害内容至关重要,但现有方法常损害生成质量。本文提出无需训练、零样本的语义手术(Semantic Surgery)框架,直接在扩散过程前操作文本嵌入。该方法动态检测提示中目标概念的存在,并进行校准向量减法以从源头中和其影响,提升擦除完整性和局部性。框架包含共现编码模块以增强多概念擦除鲁棒性,以及视觉反馈回路应对潜在概念残留。作为无训练方法,其可动态适配每条提示,实现精确干预。在物体、显性内容、艺术风格及多名人擦除任务上的大量实验表明,本方法显著优于现有最佳方案:物体擦除完成度达93.58 H-score,显性内容仅剩1例,风格擦除保持8.09 H_a且无质量下降。该鲁棒性还使框架具备内置威胁检测能力,为更安全的文本到图像生成提供实用解决方案。

原文摘要 · Abstract (English)

Concept erasure in text-to-image diffusion models is crucial for mitigating harmful content, yet existing methods often compromise generative quality. We introduce Semantic Surgery, a novel training-free, zero-shot framework for concept erasure that operates directly on text embeddings before the diffusion process. It dynamically estimates the presence of target concepts in a prompt and performs a calibrated vector subtraction to neutralize their influence at the source, enhancing both erasure completeness and locality. The framework includes a Co-Occurrence Encoding module for robust multi-concept erasure and a visual feedback loop to address latent concept persistence. As a training-free method, Semantic Surgery adapts dynamically to each prompt, ensuring precise interventions. Extensive experiments on object, explicit content, artistic style, and multi-celebrity erasure tasks show our method significantly outperforms state-of-the-art approaches. We achieve superior completeness and robustness while preserving locality and image quality (e.g., 93.58 H-score in object erasure, reducing explicit content to just 1 instance, and 8.09 H_a in style erasure with no quality degradation). This robustness also allows our framework to function as a built-in threat detection system, offering a practical solution for safer text-to-image generation.

概念擦除扩散模型安全生成零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。