arXiv:2512.13039cs.CVcs.CR2025-12

双向协同去概念化,让生成更安全且画质更好

Bi-Erasing: A Bidirectional Framework for Concept Removal in Diffusion Models

  • 用正负双分支同时抑制有害概念与引导安全图像
  • 在保持图像质量前提下,显著提升概念移除效果
  • 适合需要安全可控生成的AI绘图应用

概念擦除通过微调扩散模型以移除不想要或有害的视觉概念,已成为缓解文本到图像模型生成不当内容的主流方法。然而,现有方法多采用单向擦除策略,或抑制目标概念,或强化安全替代项,难以在概念移除与生成质量间取得平衡。为此,我们提出一种新的双向图像引导概念擦除框架(Bi-Erasing),实现有害语义抑制与安全生成增强的同步进行。基于文本提示与对应图像的联合表示,Bi-Erasing引入两个解耦的图像分支:负向分支负责抑制有害语义,正向分支则提供安全替代的视觉引导。通过联合优化这两个互补方向,该方法在擦除效果与生成可用性之间取得良好平衡。此外,我们在图像分支中引入基于掩码的过滤机制,避免无关内容干扰擦除过程。大量实验表明,所提方法在概念移除有效性与视觉保真度之间均优于基线方法。

原文摘要 · Abstract (English)

Concept erasure, which fine-tunes diffusion models to remove undesired or harmful visual concepts, has become a mainstream approach to mitigating unsafe or illegal image generation in text-to-image models.However, existing removal methods typically adopt a unidirectional erasure strategy by either suppressing the target concept or reinforcing safe alternatives, making it difficult to achieve a balanced trade-off between concept removal and generation quality. To address this limitation, we propose a novel Bidirectional Image-Guided Concept Erasure (Bi-Erasing) framework that performs concept suppression and safety enhancement simultaneously. Specifically, based on the joint representation of text prompts and corresponding images, Bi-Erasing introduces two decoupled image branches: a negative branch responsible for suppressing harmful semantics and a positive branch providing visual guidance for safe alternatives. By jointly optimizing these complementary directions, our approach achieves a balance between erasure efficacy and generation usability. In addition, we apply mask-based filtering to the image branches to prevent interference from irrelevant content during the erasure process. Across extensive experiment evaluations, the proposed Bi-Erasing outperforms baseline methods in balancing concept removal effectiveness and visual fidelity.

扩散模型概念擦除安全生成图像引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。