精准擦除图像生成模型中的特定概念,同时保留相关知识。
Fine-Grained Erasure in Text-to-Image Diffusion-based Foundation Models
- 构建概念邻域识别相关概念集,实现邻近感知的去学习。
- 在多个数据集上实现目标概念擦除,相关概念保留率提升至少12%。
- 适合需要可控内容删除的AI安全与隐私保护场景。
现有文本到图像生成模型中的去学习算法在移除特定目标概念时,常导致语义相关概念的知识丢失,这一问题称为邻接性挑战。为解决该问题,我们提出FADE(Fine-grained Attenuation for Diffusion Erasure),引入扩散模型中的邻接感知去学习机制。FADE包含两个组件:(1) 概念邻域,用于识别相关概念集合;(2) 网格模块,通过结构化组合清除、邻接与引导损失,实现对目标概念的精确擦除,同时保持相关与无关概念的生成保真度。在Stanford Dogs、Oxford Flowers、CUB、I2P、Imagenette和ImageNet1k等数据集上评估,FADE有效移除目标概念,对相关概念影响极小,在保留性能上相较当前最优方法至少提升12%。
原文摘要 · Abstract (English)
Existing unlearning algorithms in text-to-image generative models often fail to preserve the knowledge of semantically related concepts when removing specific target concepts: a challenge known as adjacency. To address this, we propose FADE (Fine grained Attenuation for Diffusion Erasure), introducing adjacency aware unlearning in diffusion models. FADE comprises two components: (1) the Concept Neighborhood, which identifies an adjacency set of related concepts, and (2) Mesh Modules, employing a structured combination of Expungement, Adjacency, and Guidance loss components. These enable precise erasure of target concepts while preserving fidelity across related and unrelated concepts. Evaluated on datasets like Stanford Dogs, Oxford Flowers, CUB, I2P, Imagenette, and ImageNet1k, FADE effectively removes target concepts with minimal impact on correlated concepts, achieving atleast a 12% improvement in retention performance over state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。