arXiv:2603.10445cs.LGcs.CV2026-03被引 1

无需提示词即可删除扩散模型中的特定内容,保护隐私与伦理

Unlearning the Unpromptable: Prompt-free Instance Unlearning in Diffusion Models

  • 用图像编辑和时间步加权实现无提示词的实例遗忘
  • 在稳定扩散3和DDPM-CelebA上成功移除人脸等不可提示内容
  • 适合需快速修复模型隐私或伦理问题的厂商使用

机器遗忘旨在从训练好的模型中移除特定输出,通常以概念层面进行,例如忘记某位名人的所有出现或通过文本提示过滤内容。然而,许多不期望的输出(如个人面部或文化/事实错误生成)往往无法通过文本提示精确描述。本文针对这一未充分探索的‘实例遗忘’场景——即目标输出不可提示但需删除的情况——提出一种基于代理的高效遗忘方法。该方法结合图像编辑、时间步感知加权与梯度手术,引导已训练的扩散模型选择性遗忘特定输出。在条件生成(Stable Diffusion 3)与无条件生成(DDPM-CelebA)扩散模型上的实验表明,该无提示方法能有效遗忘不可提示的内容(如人脸、文化误判图像),同时保持其余生成质量不受损,优于现有提示型与无提示基线。该方法可为扩散模型提供实用的隐私与伦理合规热修复方案。

原文摘要 · Abstract (English)

Machine unlearning aims to remove specific outputs from trained models, often at the concept level, such as forgetting all occurrences of a particular celebrity or filtering content via text prompts. However, many undesired outputs, such as an individual's face or generations culturally or factually misinterpreted, cannot often be specified by text prompts. We address this underexplored setting of instance unlearning for outputs that are undesired but unpromptable, where the goal is to forget target outputs selectively while preserving the rest. To this end, we introduce an effective surrogate-based unlearning method that leverages image editing, timestep-aware weighting, and gradient surgery to guide trained diffusion models toward forgetting specific outputs. Experiments on conditional (Stable Diffusion 3) and unconditional (DDPM-CelebA) diffusion models demonstrate that our prompt-free method uniquely unlearns unpromptable outputs, such as faces and culturally inaccurate depictions, with preserved integrity, unlike prompt-based and prompt-free baselines. Our proposed method would serve as a practical hotfix for diffusion model providers to ensure privacy protection and ethical compliance.

扩散模型实例遗忘隐私保护无提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。