arXiv:2605.12122cs.LGcs.AI2026-05被引 1

提出新方法实现扩散模型中概念的精准清除,减少误伤。

Disentangled Sparse Representations for Concept-Separated Diffusion Unlearning

论文配图:Disentangled Sparse Representations for Concept-Separated Diffusion Unlearning
图 1 · 摘自论文原文
  • 用对比学习让潜空间按概念分组,避免特征混杂。
  • 在联合风格-物体删除任务上效果领先,干扰更小。
  • 适合需要精确控制生成内容的场景,如安全合规。

在文本到图像扩散模型中消除特定概念的重要性日益凸显,以防止不当内容生成。现有基于稀疏自编码器(SAE)的方法通过轻量级调整潜在特征实现概念抑制,无需修改模型参数。然而,仅依赖稀疏重建目标训练的SAE无法显式保证概念间分离,导致不同概念共享潜在特征。为此,我们提出SAEParate,通过概念感知的对比目标将潜在表示组织为概念专属聚类,实现更精准的概念抑制,同时降低未删概念的意外干扰。此外,我们采用基于GeLU的非线性变换增强编码器表达能力,在此分离目标下构建更具判别性的解耦潜在空间。在UnlearnCanvas上的实验表明,该方法达到当前最优性能,尤其在联合风格-物体删除这一高难度场景中表现突出,显著缓解了目标与非目标概念间的严重干扰。

原文摘要 · Abstract (English)

Unlearning specific concepts in text-to-image diffusion models has become increasingly important for preventing undesirable content generation. Among prior approaches, sparse autoencoder (SAE)-based methods have attracted attention due to their ability to suppress target concepts through lightweight manipulation of latent features, without modifying model parameters. However, SAEs trained with sparse reconstruction objectives do not explicitly enforce concept-wise separation, resulting in shared latent features across concepts. To address this, we propose SAEParate, which organizes latent representations into concept-specific clusters via a concept-aware contrastive objective, enabling more precise concept suppression while reducing unintended interference during unlearning. In addition, we enhance the encoder with a GeLU-based nonlinear transformation to increase its expressive capacity under this separation objective, enabling a more discriminative and disentangled latent space. Experiments on UnlearnCanvas demonstrate state-of-the-art performance, with particularly strong gains in joint style-object unlearning, a challenging setting where existing methods suffer from severe interference between target and non-target concepts.

扩散模型去概念化解耦表示生成安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。