让图像生成模型忘记特定概念,同时保持其他生成质量与多样性。
TILDE: TILt-based Distributional Erasure for Concept Unlearning

- 通过分布对齐方法,精准移除目标概念而不破坏模型整体性能。
- 在物体、风格和角色上均实现强遗忘效果,且生成质量优于现有方法。
- 适合需安全合规的生成模型部署,如版权规避或隐私保护场景。
文本到图像扩散模型中的概念删除对安全与实际应用至关重要:随着隐私担忧、版权纠纷、商标限制及安全法规上升,已部署系统必须能在训练后抑制不希望存在的概念。现有方法虽能有效移除目标概念,但实际删除还应具备关键特性——删除后的模型需保留生成质量、多样性和语义覆盖范围。理想状态是仅用良性数据从头训练的模型。然而,常见删除目标未明确应逼近何种后续分布,导致保留性成为更新规则的隐含结果。本文提出TILDE(基于倾斜的分布擦除),将概念删除建模为分布对齐问题:目标是满足遗忘约束下,预训练模型条件分布中偏差最小的分布。该能量倾斜、无锚点的目标可抑制含概念的图像,同时保留各提示下的良性相对概率质量。我们通过残差∇-GFlowNet训练实现该原则,学习相对于预训练扩散模型的遗忘能量所诱导的得分修正。在物体、艺术风格和角色上,TILDE均实现了强遗忘效果,并在保留性和分布保真度方面超越已有基线。
原文摘要 · Abstract (English)
Concept unlearning in text-to-image diffusion models is critical for safe and practical deployment: with rising privacy concerns, copyright disputes, trademark constraints, and safety regulations, deployed systems must be able to suppress unwanted concepts after training. Existing methods often remove the target concept effectively, but practical unlearning also requires an equally fundamental property: the unlearned model should retain quality, diversity, and semantic coverage on benign generation. The gold standard is a retain-only model trained from scratch without the unwanted data. However, common erasure objectives do not specify which post-unlearning distribution should approximate this reference, leaving retention as an implicit consequence of the update rule. We propose TILDE, TILt-based Distributional Erasure, which formulates concept unlearning as a distributional alignment problem: the desired target is the minimum-deviation conditional distribution from the pretrained model under a forgetting constraint. This energy-tilted, anchor-free target suppresses concept-expressing images while preserving benign relative mass for each prompt. We instantiate this principle with residual $\nabla$-GFlowNet training, which learns the score correction induced by the forget energy relative to the pretrained diffusion model. Across objects, artistic styles, and characters, TILDE achieves strong forgetting while improving retention and distributional fidelity over prior baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。