arXiv:2603.18767cs.AI2026-03被引 1

用多样提示词精准删除图像模型中的有害概念

A Concept is More Than a Word: Diversified Unlearning in Text-to-Image Diffusion Models

  • 用一组上下文多样的提示词代替单个关键词表示概念
  • 相比传统方法,删得更准且不误伤其他内容
  • 适合需要安全可控生成的AI应用开发者

概念消去已成为减少文本到图像扩散模型生成有害内容风险的有前景方向,通过选择性地从模型参数中擦除不良概念。现有方法通常依赖关键词来识别需消去的目标概念。然而我们发现,这种基于关键词的设定存在根本局限:一个视觉概念具有多维特征,可通过多种文本形式表达,且在潜在空间中常与相关概念重叠,导致仅靠关键词的消去方式不够精确,容易过度遗忘。这是因为单一关键词仅代表概念的一个狭窄点估计,无法覆盖其完整语义分布及潜在空间中的复杂变化。为解决此问题,我们提出多样化消去(Diversified Unlearning),一种通过一组上下文多样的提示词而非单个关键词来表征概念的分布式框架。该更丰富的表示实现了更精准、更鲁棒的消去效果。在多个基准和先进基线上的大量实验表明,将多样化消去作为附加组件集成到现有消去流程中,能持续实现更强的擦除效果、更好的无关概念保留能力,并提升对抗性恢复攻击下的鲁棒性。

原文摘要 · Abstract (English)

Concept unlearning has emerged as a promising direction for reducing the risks of harmful content generation in text-to-image diffusion models by selectively erasing undesirable concepts from a model's parameters. Existing approaches typically rely on keywords to identify the target concept to be unlearned. However, we show that this keyword-based formulation is inherently limited: a visual concept is multi-dimensional, can be expressed in diverse textual forms, and often overlap with related concepts in the latent space, making keyword-only unlearning, which imprecisely indicate the target concept is brittle and prone to over-forgetting. This occurs because a single keyword represents only a narrow point estimate of the concept, failing to cover its full semantic distribution and entangled variations in the latent space. To address this limitation, we propose Diversified Unlearning, a distributional framework that represents a concept through a set of contextually diverse prompts rather than a single keyword. This richer representation enables more precise and robust unlearning. Through extensive experiments across multiple benchmarks and state-of-the-art baselines, we demonstrate that integrating Diversified Unlearning as an add-on component into existing unlearning pipelines consistently achieves stronger erasure, better retention of unrelated concepts, and improved robustness against adversarial recovery attacks.

概念消去扩散模型文本生成图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。