arXiv:2503.13769cs.CV2025-03被引 10

无需重训即可逐步移除生成模型中的特定概念,且不损失其他能力。

Continual Unlearning for Foundational Text-to-Image Models without Generalization Erosion

  • 提出增量式持续遗忘方法,分步移除目标概念。
  • 在移除指定概念的同时,保持非目标概念的生成能力。
  • 适合需清理版权或风格侵权内容的生成模型优化场景。

如何在不进行大规模重训练的情况下,有效从预训练生成模型中移除特定概念?本研究提出「持续遗忘」新范式,实现对基础生成模型中多个特定概念的渐进式靶向清除。我们提出无泛化退化的减量遗忘(DUGE)算法,可选择性地消除目标概念的生成,同时保留相关非目标概念的生成能力,并缓解泛化性能下降问题。DUGE通过三项损失实现:交叉注意力损失引导模型关注不含目标概念的图像;先验保持损失保护非目标概念知识;正则化损失防止泛化能力退化。实验表明,该方法可在不损害模型整体完整性与性能的前提下,排除特定概念。这为生成模型的精细化调整提供了实用方案,能有效应对版权侵权、个人或受版权材料滥用、独特艺术风格复制等风险。重要的是,非目标概念得以保留,保障了模型核心能力与有效性。

原文摘要 · Abstract (English)

How can we effectively unlearn selected concepts from pre-trained generative foundation models without resorting to extensive retraining? This research introduces `continual unlearning', a novel paradigm that enables the targeted removal of multiple specific concepts from foundational generative models, incrementally. We propose Decremental Unlearning without Generalization Erosion (DUGE) algorithm which selectively unlearns the generation of undesired concepts while preserving the generation of related, non-targeted concepts and alleviating generalization erosion. For this, DUGE targets three losses: a cross-attention loss that steers the focus towards images devoid of the target concept; a prior-preservation loss that safeguards knowledge related to non-target concepts; and a regularization loss that prevents the model from suffering from generalization erosion. Experimental results demonstrate the ability of the proposed approach to exclude certain concepts without compromising the overall integrity and performance of the model. This offers a pragmatic solution for refining generative models, adeptly handling the intricacies of model training and concept management lowering the risks of copyright infringement, personal or licensed material misuse, and replication of distinctive artistic styles. Importantly, it maintains the non-targeted concepts, thereby safeguarding the model's core capabilities and effectiveness.

生成模型概念遗忘模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。