arXiv:2602.19631cs.CVcs.AI2026-02中稿 · ICLR被引 2

通过误导文本编码器高层语义,精准擦除图像生成中的特定概念。

Localized Concept Erasure in Text-to-Image Diffusion Models via High-Level Representation Misdirection

  • 在文本编码器早期层误导目标概念的高层语义表示。
  • 在多个数据集上实现强去概念效果,且不影响其他内容生成质量。
  • 低训练成本,兼容主流模型,适合需可控生成的场景。

近期文本到图像扩散模型快速发展并广泛应用,但其强大的生成能力也带来生成有害、隐私或版权内容的风险。为此,概念擦除技术成为有前景的解决方案。以往方法主要针对去噪模块(如U-Net主干)进行微调,但最新因果追踪研究指出,视觉属性信息集中于文本编码器的早期自注意力层,提示了新的擦除路径。我们初步实验发现直接微调早期层虽能抑制目标概念,但常损害非目标概念的生成质量。为此,我们提出高层表示误导(HiRM),将目标概念的高层语义表示引导至随机方向或语义定义方向(如超类别),仅更新包含因果状态的早期层。该解耦策略实现精准概念移除,对无关概念影响极小,在UnlearnCanvas和NSFW基准上对多种目标(如物体、风格、裸露)均表现优异。HiRM保持生成实用性,训练成本低,可迁移至Flux等先进架构而无需额外训练,并与基于去噪器的方法产生协同效应。

原文摘要 · Abstract (English)

Recent advances in text-to-image (T2I) diffusion models have seen rapid and widespread adoption. However, their powerful generative capabilities raise concerns about potential misuse for synthesizing harmful, private, or copyrighted content. To mitigate such risks, concept erasure techniques have emerged as a promising solution. Prior works have primarily focused on fine-tuning the denoising component (e.g., the U-Net backbone). However, recent causal tracing studies suggest that visual attribute information is localized in the early self-attention layers of the text encoder, indicating a potential alternative for concept erasing. Building on this insight, we conduct preliminary experiments and find that directly fine-tuning early layers can suppress target concepts but often degrades the generation quality of non-target concepts. To overcome this limitation, we propose High-Level Representation Misdirection (HiRM), which misdirects high-level semantic representations of target concepts in the text encoder toward designated vectors such as random directions or semantically defined directions (e.g., supercategories), while updating only early layers that contain causal states of visual attributes. Our decoupling strategy enables precise concept removal with minimal impact on unrelated concepts, as demonstrated by strong results on UnlearnCanvas and NSFW benchmarks across diverse targets (e.g., objects, styles, nudity). HiRM also preserves generative utility at low training cost, transfers to state-of-the-art architectures such as Flux without additional training, and shows synergistic effects with denoiser-based concept erasing methods.

概念擦除扩散模型文本编码器可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。