arXiv:2510.03302cs.LGcs.CV2025-10被引 14

用强化学习让被删除的概念在扩散模型中复活,揭示现有安全机制的漏洞。

Revoking Amnesia: RL-based Trajectory Optimization to Resurrect Erased Concepts in Diffusion Models

  • 通过强化学习动态调整去噪过程,不修改模型权重实现概念复活。
  • 在多个数据集上恢复效果优于现有方法,计算时间减少10倍。
  • 适合关注生成模型安全机制可靠性的研究人员和开发者。

概念擦除技术广泛应用于文本到图像扩散模型中,以防止不当内容生成,保障安全与版权。然而,随着模型演进至Flux等下一代架构,现有擦除方法(如ESD、UCE、AC)的效果显著下降,引发对其真实机制的质疑。系统性分析表明,概念擦除仅制造了‘遗忘’的假象:并非真正遗忘,而是使采样轨迹偏离目标概念,因此擦除本质上可逆。这一发现促使我们区分表面安全与真正的概念移除。本文提出 extbf{RevAm}(​Revoking ​Amnesia),一种基于强化学习的轨迹优化框架,通过动态引导去噪过程,在不修改模型权重的情况下复活被擦除概念。通过将群体相对策略优化(GRPO)适配至扩散模型,RevAm利用轨迹级奖励探索多样化恢复路径,克服了现有方法受限于局部最优的问题。大量实验表明,RevAm在概念复活保真度上表现更优,同时计算时间降低10倍,暴露出当前安全机制的关键脆弱性,并强调需要超越轨迹操控的更稳健擦除技术。

原文摘要 · Abstract (English)

Concept erasure techniques have been widely deployed in T2I diffusion models to prevent inappropriate content generation for safety and copyright considerations. However, as models evolve to next-generation architectures like Flux, established erasure methods (\textit{e.g.}, ESD, UCE, AC) exhibit degraded effectiveness, raising questions about their true mechanisms. Through systematic analysis, we reveal that concept erasure creates only an illusion of ``amnesia": rather than genuine forgetting, these methods bias sampling trajectories away from target concepts, making the erasure fundamentally reversible. This insight motivates the need to distinguish superficial safety from genuine concept removal. In this work, we propose \textbf{RevAm} (\underline{Rev}oking \underline{Am}nesia), an RL-based trajectory optimization framework that resurrects erased concepts by dynamically steering the denoising process without modifying model weights. By adapting Group Relative Policy Optimization (GRPO) to diffusion models, RevAm explores diverse recovery trajectories through trajectory-level rewards, overcoming local optima that limit existing methods. Extensive experiments demonstrate that RevAm achieves superior concept resurrection fidelity while reducing computational time by 10$\times$, exposing critical vulnerabilities in current safety mechanisms and underscoring the need for more robust erasure techniques beyond trajectory manipulation.

扩散模型概念擦除强化学习安全机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。