提出新方法保护扩散模型中被误删的关联概念。
Co-occurring Associated REtained concepts in Diffusion Unlearning

- 构建CARE集,自动提取需保留的关联词元用于训练。
- 在去除目标概念时,保持其他良性概念不被抑制。
- 适用于图像生成中避免误删人物等关键元素的场景。
去学习已成为缓解扩散模型生成有害内容的关键技术。然而,现有方法常不仅移除目标概念,还会误删良性共现概念。如图1所示,删除裸露概念会意外抑制人物概念,导致模型无法生成含人物的图像。本文定义这些必须保留的非目标共现概念为CARE(Co-occurring Associated REtained concepts),并提出CARE分数作为衡量其保留程度的通用指标。基于此,提出ReCARE框架,显式保护CARE同时仅擦除目标概念。ReCARE自动构建CARE集,从目标图像中提取良性共现词元,并在训练中利用该词汇表实现稳定去学习。在多种目标概念(裸露、梵高风格、天竺鲷对象)上的大量实验表明,ReCARE在鲁棒概念擦除、整体性能与CARE保留之间实现了当前最优平衡。
原文摘要 · Abstract (English)
Unlearning has emerged as a key technique to mitigate harmful content generation in diffusion models. However, existing methods often remove not only the target concept, but also benign co-occurring concepts. As illustrated in Fig.1, unlearning nudity can unintentionally suppress the concept of person, preventing a model from generating images with person. We define these undesirably suppressed co-occurring concepts that must be preserved CARE (Co-occurring Associated REtained concepts). Then, we introduce the CARE score, a general metric that directly quantifies their preservation across unlearning tasks. With this foundation, we propose ReCARE (Robust erasure for CARE), a framework that explicitly safeguards CARE while erasing only the target concept. ReCARE automatically constructs the CARE-set, a curated vocabulary of benign co-occurring tokens extracted from target images, and leverages this vocabulary during training for stable unlearning. Extensive experiments across various target concepts (Nudity, Van Gogh style, and Tench object) demonstrate that ReCARE achieves overall state-of-the-art performance in balancing robust concept erasure, overall utility, and CARE preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。