arXiv:2604.16483cs.CVcs.AI2026-04

提出动态擦除框架,精准删除图像生成中的敏感概念而不失真

Dynamic Eraser for Guided Concept Erasure in Diffusion Models

论文配图:Dynamic Eraser for Guided Concept Erasure in Diffusion Models
图 1 · 摘自论文原文
  • 通过自动识别安全语义锚点,实现对敏感内容的精准定位
  • 91%的敏感概念擦除率,显著优于现有方法(最高85.9%)
  • 无需训练、轻量级,适合需要可控内容生成的场景

文本到图像扩散模型中的概念擦除对安全内容生成至关重要,但现有推理阶段方法存在明显局限。特征修正类方法常导致过度纠正,而基于标记的干预难以兼顾语义粒度与上下文一致性,且易引发严重语义漂移甚至表征崩溃。为此,我们提出轻量、免训练的动态语义引导(DSS)框架,实现可解释、可控制的概念擦除。DSS引入:1)敏感语义边界建模(SSBM),自动化发现安全语义锚点;2)敏感语义引导(SSG),利用交叉注意力特征精确检测并基于良好定义目标的闭式解进行修正。该方法在保证敏感内容有效抑制的同时,最大限度保留良性语义。实验表明,DSS平均擦除率达91.0%,显著超越当前最优方法(18.6%~85.9%),且对输出保真度影响极小。

原文摘要 · Abstract (English)

Concept erasure in Text-To-Image (T2I) diffusion models is vital for safe content generation, but existing inference-time methods face significant limitations. Feature-correction approaches often cause uncontrolled over-correction, while token-level interventions struggle with semantic granularity and context. Moreover, both types of methods are prone to severe semantic drift or even complete representation collapse. To address these challenges, we present Dynamic Semantic Steering (DSS), a lightweight, training-free framework for interpretable and controllable concept erasure. DSS introduces: 1) Sensitive Semantic Boundary Modeling (SSBM) to automate the discovery of safe semantic anchors, and 2) Sensitive Semantic Guidance (SSG), which leverages cross-attention features for precise detection and performs correction via a closed-form solution derived from a well-posed objective. This ensures optimal suppression of sensitive content while preserving benign semantics. DSS achieves an average erasure rate of 91.0\%, significantly outperforming SOTA methods (from 18.6\% to 85.9\%) with minimal impact on output fidelity.

扩散模型概念擦除安全生成语义控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。