arXiv:2507.12283cs.CV2025-07

让扩散模型删除特定敏感概念,保护隐私与公平性

FADE: Adversarial Concept Erasure in Flow Models

  • 通过对抗目标与轨迹感知微调,精准擦除指定概念
  • 在多个任务上实现领先去除效果,保真度提升5%-10%
  • 适合需要可控内容生成的AI安全与伦理研究者

扩散模型虽具备强大图像生成能力,但可能记忆敏感信息或延续偏见。本文提出一种新的概念擦除方法FADE(Fair Adversarial Diffusion Erasure),旨在从文本到图像的扩散模型中移除特定概念(如个人身份、有害刻板印象)。该方法结合轨迹感知微调与对抗目标,确保概念被可靠清除的同时保持模型整体生成质量。理论上,证明了该方法可最小化被擦除概念与模型输出间的互信息,保障隐私与公平。实验在Stable Diffusion和FLUX上评估,使用MACE基准中的物体、名人、成人内容及风格擦除任务,结果表明FADE在去除效果和图像质量上超越ESD、UCE、MACE和ANT等现有基线,其概念去除与保真度的调和均值提升5%–10%。消融实验证实对抗目标与轨迹保持机制均对性能有贡献。本工作为安全、公平的生成建模设立了新标准,无需从头训练即可消除指定概念。

原文摘要 · Abstract (English)

Diffusion models have demonstrated remarkable image generation capabilities, but also pose risks in privacy and fairness by memorizing sensitive concepts or perpetuating biases. We propose a novel \textbf{concept erasure} method for text-to-image diffusion models, designed to remove specified concepts (e.g., a private individual or a harmful stereotype) from the model's generative repertoire. Our method, termed \textbf{FADE} (Fair Adversarial Diffusion Erasure), combines a trajectory-aware fine-tuning strategy with an adversarial objective to ensure the concept is reliably removed while preserving overall model fidelity. Theoretically, we prove a formal guarantee that our approach minimizes the mutual information between the erased concept and the model's outputs, ensuring privacy and fairness. Empirically, we evaluate FADE on Stable Diffusion and FLUX, using benchmarks from prior work (e.g., object, celebrity, explicit content, and style erasure tasks from MACE). FADE achieves state-of-the-art concept removal performance, surpassing recent baselines like ESD, UCE, MACE, and ANT in terms of removal efficacy and image quality. Notably, FADE improves the harmonic mean of concept removal and fidelity by 5--10\% over the best prior method. We also conduct an ablation study to validate each component of FADE, confirming that our adversarial and trajectory-preserving objectives each contribute to its superior performance. Our work sets a new standard for safe and fair generative modeling by unlearning specified concepts without retraining from scratch.

扩散模型概念擦除隐私保护公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。