arXiv:2511.00804cs.LGcs.CV2025-11NeurIPS被引 9

用轨迹探索方法实现图像生成中的概念擦除,不降质且泛化强。

EraseFlow: Learning Concept Erasure Policies via GFlowNet-Driven Alignment

  • 通过GFlowNet采样完整去噪路径,学习概念擦除策略
  • 在不依赖奖励模型情况下,实现高质量图像生成与概念消除
  • 适合需要安全可控生成的AI系统开发者使用

从强大的文本到图像生成器中擦除有害或专有概念是一项新兴的安全需求。现有概念擦除技术要么导致图像质量下降,依赖脆弱的对抗性损失,要么需要代价高昂的再训练周期。我们发现这些局限源于对扩散生成中去噪轨迹的狭隘理解。本文提出EraseFlow,首个将概念遗忘建模为去噪路径空间探索的框架,并采用具有轨迹平衡目标的GFlowNets进行优化。通过采样整个轨迹而非单个终点状态,EraseFlow学习到一种随机策略,能在避免目标概念的同时保持模型先验。该方法无需精心设计的奖励模型,因此能有效泛化至未见概念,避免可被破解的奖励机制,同时提升性能。大量实验证明,EraseFlow优于现有基线,在性能与先验保留之间取得最佳平衡。

原文摘要 · Abstract (English)

Erasing harmful or proprietary concepts from powerful text to image generators is an emerging safety requirement, yet current "concept erasure" techniques either collapse image quality, rely on brittle adversarial losses, or demand prohibitive retraining cycles. We trace these limitations to a myopic view of the denoising trajectories that govern diffusion based generation. We introduce EraseFlow, the first framework that casts concept unlearning as exploration in the space of denoising paths and optimizes it with GFlowNets equipped with the trajectory balance objective. By sampling entire trajectories rather than single end states, EraseFlow learns a stochastic policy that steers generation away from target concepts while preserving the model's prior. EraseFlow eliminates the need for carefully crafted reward models and by doing this, it generalizes effectively to unseen concepts and avoids hackable rewards while improving the performance. Extensive empirical results demonstrate that EraseFlow outperforms existing baselines and achieves an optimal trade off between performance and prior preservation.

概念擦除扩散模型生成安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。