无需训练即可精准擦除扩散模型中的特定概念,效果优于现有方法。
ActErase: A Training-Free Paradigm for Precise Concept Erasure via Activation Redirection
- 通过提示对分析激活差异区域,动态替换输入激活实现概念擦除。
- 在三类擦除任务中均达当前最优性能,且生成能力保持稳定。
- 无需微调、抗攻击性强,适合快速部署到各类扩散模型中。
文本到图像扩散模型虽具备强大生成能力,但引发安全、版权与伦理问题。现有概念擦除方法多依赖数据密集型且计算成本高的微调,存在显著局限。受启发于模型激活主要由通用概念构成,目标概念仅占极小部分,我们提出一种全新的无训练方法(ActErase),通过提示对分析识别激活差异区域,提取目标激活并在前向传播中动态替换输入激活。在三类关键擦除任务(裸露内容、艺术风格、物体移除)上的综合评估表明,该方法达到当前最优擦除性能,同时有效保留模型整体生成能力。此外,方法对对抗攻击具有强鲁棒性,建立了一种轻量级且高效的即插即用型概念操控范式。
原文摘要 · Abstract (English)
Recent advances in text-to-image diffusion models have demonstrated remarkable generation capabilities, yet they raise significant concerns regarding safety, copyright, and ethical implications. Existing concept erasure methods address these risks by removing sensitive concepts from pre-trained models, but most of them rely on data-intensive and computationally expensive fine-tuning, which poses a critical limitation. To overcome these challenges, inspired by the observation that the model's activations are predominantly composed of generic concepts, with only a minimal component can represent the target concept, we propose a novel training-free method (ActErase) for efficient concept erasure. Specifically, the proposed method operates by identifying activation difference regions via prompt-pair analysis, extracting target activations and dynamically replacing input activations during forward passes. Comprehensive evaluations across three critical erasure tasks (nudity, artistic style, and object removal) demonstrates that our training-free method achieves state-of-the-art (SOTA) erasure performance, while effectively preserving the model's overall generative capability. Our approach also exhibits strong robustness against adversarial attacks, establishing a new plug-and-play paradigm for lightweight yet effective concept manipulation in diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。