通过安全-不安全配对实现概念擦除,保持生成图像的一致性。
Consistency-Preserving Concept Erasure via Unsafe-Safe Pairing and Directional Fisher-weighted Adaptation
- 用不安全输入生成对应安全图像,形成成对数据提升语义一致性。
- 在多个数据集上比现有方法减少90%以上语义失真,保留高质量生成。
- 适合需要精准删除有害内容又不破坏图像结构的场景。
随着文本到图像扩散模型的日益多功能化,选择性擦除不良概念(如有害内容)变得至关重要。然而,现有方法主要关注移除不安全概念,缺乏对相应安全替代物的引导,常导致原始生成与擦除后生成在结构和语义上不一致。本文提出新框架PAIRed Erasing(PAIR),将概念擦除从简单删除重构为一致性保持的语义重对齐,利用不安全-安全配对数据。首先生成与不安全输入对应的结构与语义一致的安全图像,构建成对的多模态数据。基于这些配对,引入两个核心组件:(1) 配对语义重对齐,使用不安全-安全配对显式地将目标概念映射至语义对齐的安全锚点;(2) Fisher加权的DoRA初始化,利用不安全-安全配对初始化参数高效低秩适配矩阵,促进生成安全替代物的同时选择性抑制不安全概念。二者协同实现细粒度擦除,仅移除目标概念,同时保持整体语义一致性。大量实验表明,本方法显著优于当前最优基线,在有效擦除概念的同时,保持结构完整性、语义连贯性和生成质量。
原文摘要 · Abstract (English)
With the increasing versatility of text-to-image diffusion models, the ability to selectively erase undesirable concepts (e.g., harmful content) has become indispensable. However, existing concept erasure approaches primarily focus on removing unsafe concepts without providing guidance toward corresponding safe alternatives, which often leads to failure in preserving the structural and semantic consistency between the original and erased generations. In this paper, we propose a novel framework, PAIRed Erasing (PAIR), which reframes concept erasure from simple removal to consistency-preserving semantic realignment using unsafe-safe pairs. We first generate safe counterparts from unsafe inputs while preserving structural and semantic fidelity, forming paired unsafe-safe multimodal data. Leveraging these pairs, we introduce two key components: (1) Paired Semantic Realignment, a guided objective that uses unsafe-safe pairs to explicitly map target concepts to semantically aligned safe anchors; and (2) Fisher-weighted Initialization for DoRA, which initializes parameter-efficient low-rank adaptation matrices using unsafe-safe pairs, encouraging the generation of safe alternatives while selectively suppressing unsafe concepts. Together, these components enable fine-grained erasure that removes only the targeted concepts while maintaining overall semantic consistency. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art baselines, achieving effective concept erasure while preserving structural integrity, semantic coherence, and generation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。