无需重训练,用能量引导实现图像生成中概念擦除
EGLOCE: Training-Free Energy-Guided Latent Optimization for Concept Erasure
- 通过潜在空间梯度下降,用排斥能量驱散目标概念
- 保留语义对齐的保持能量,确保提示一致性与图像质量
- 完全在推理阶段操作,可即插即用,抗对抗攻击
随着文本到图像扩散模型日益普及,移除特定概念(如敏感内容、受版权保护的角色或风格)已成为保障安全与合规的必要需求。现有遗忘方法常需代价高昂的再训练,或修改模型参数导致无关概念保真度下降,亦或依赖间接的推理时调整,削弱概念擦除效果。受扩散模型条件保持中能量引导采样的启发,本文提出训练自由的概念擦除方法EGLOCE,通过在推理阶段重新定向噪声潜在变量来移除不需要的概念。该方法采用双目标框架:排斥能量通过潜在空间梯度下降使生成远离目标概念,保持能量则维护原始提示的语义对齐。相比需修改模型权重或提供弱推理引导的先前方法,EGLOCE完全在推理阶段运行,提升擦除性能,并支持即插即用集成。大量实验表明,即便面对对抗攻击,EGLOCE仍能有效移除概念并维持图像质量和提示一致性。据我们所知,本工作首次建立了一种基于采样过程中双能量引导的安全可控图像生成新范式。
原文摘要 · Abstract (English)
As text-to-image diffusion models grow increasingly prevalent, the ability to remove specific concepts-mostly explicit content and many copyrighted characters or styles-has become essential for safety and compliance. Existing unlearning approaches often require costly re-training, modify parameters at the cost of degradation of unrelated concept fidelity, or depend on indirect inference-time adjustment that compromise the effectiveness of concept erasure. Inspired by the success of energy-guided sampling for preservation of the condition of diffusion models, we introduce Energy-Guided Latent Optimization for Concept Erasure (EGLOCE), a training-free approach that removes unwanted concepts by re-directing noisy latent during inference. Our method employs a dual-objective framework: a repulsion energy that steers generation away from target concepts via gradient descent in latent space, and a retention energy that preserves semantic alignment to the original prompt. Combined with previous approaches that either require erroneous modified model weights or provide weak inference-time guidance, EGLOCE operates entirely at inference and enhances erasure performance, enabling plug-and-play integration. Extensive experiments demonstrate that EGLOCE improves concept removal while maintaining image quality and prompt alignment across baselines, even with adversarial attacks. To the best of our knowledge, our work is the first to establish a new paradigm for safe and controllable image generation through dual energy-based guidance during sampling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。