arXiv:2608.07620cs.CV2026-08

通过在线策略蒸馏实现多概念消除,提升生成安全性和可控性。

FlowErase-OPD: Multi-Concept Erasure via Anchored On-Policy Distillation in Flow Matching Models

论文配图:FlowErase-OPD: Multi-Concept Erasure via Anchored On-Policy Distillation in Flow Matching Models
图 1 · 摘自论文原文
  • 基于在线策略蒸馏整合多个去概念模型,统一生成控制。
  • 在去除多类有害内容时保持图像质量与语义一致性,效果优于现有方法。
  • 适合关注生成安全、内容可控的AI图像生成研究者使用。

流匹配模型在文本到图像生成中显著提升了质量,但也引发了生成有害或不期望内容的安全担忧。现有概念消除方法多针对单一概念,同时消除多个概念仍具挑战。本文提出FlowErase-OPD,一种基于在线策略蒸馏(OPD)的多概念消除框架。首先将多个单概念消除模型蒸馏为统一的LoRA模块,并引入锚定多教师蒸馏(AMTD),通过保留教师缓解消除与生成能力之间的权衡。为更好协调多目标消除,设计自适应保留控制(ARC),动态调节各消除教师的采样频率、损失权重及消除与保留教师的相对贡献。在裸露、物体和艺术风格等多类消除任务上的实验表明,FlowErase-OPD在消除有效性、图像质量和语义对齐间取得更优平衡,性能达当前最优。此外,生成模型对对抗攻击表现出强鲁棒性。结果表明,在线策略蒸馏是流匹配模型中实现安全可控生成的可行范式。

原文摘要 · Abstract (English)

Recent advances in flow matching models have substantially improved the quality of text-to-image generation, but have also raised increasing safety concerns due to their potential to generate harmful or undesirable content. Existing concept erasure methods for flow matching models predominantly focus on removing individual concepts, while effectively erasing multiple concepts simultaneously remains challenging. We propose FlowErase-OPD, a framework for multi-concept erasure based on on-policy distillation (OPD). Our approach first distills multiple single-concept erased models into a unified LoRA module and introduces Anchored Multi-Teacher Distillation (AMTD), which incorporates a retention teacher to mitigate the trade-off between concept erasure and preservation of generative capabilities. To further improve the coordination of multiple erasure objectives, we develop Adaptive Retention Control (ARC), which dynamically adjusts the sampling frequency and loss weight of each erasure teacher, together with the relative contribution of erasure and retention teachers throughout training. Extensive experiments on nudity, object, and artistic-style erasure demonstrate that FlowErase-OPD consistently improves the trade-off between erasure effectiveness, image quality, and semantic alignment, achieving state-of-the-art performance across diverse multi-concept erasure settings. Furthermore, the resulting models exhibit strong robustness against adversarial attacks. These results highlight the potential of on-policy distillation as a principled framework for safe and controllable generation in flow matching models.

图像生成安全控制流匹配多概念消除

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。