arXiv:2507.13386cs.CVcs.LG2025-07ICML被引 16

提出最小化修改的图像生成概念擦除方法,不降性能也能去除非期望内容。

Minimalist Concept Erasure in Generative Models

  • 仅通过生成结果分布距离设计目标,减少模型改动。
  • 在流匹配模型上实现概念擦除且不影响整体生成质量。
  • 适合关注生成模型安全与版权的开发者和研究者使用。

生成模型虽能产出高质量图像,但依赖大规模无标注数据引发安全与版权问题。现有概念擦除方法常过度修改模型,损害整体性能。本文提出一种基于生成输出分布距离的最小化概念擦除目标,构建可微优化损失,支持端到端反向传播。通过理论分析揭示与现有方法的联系,并引入神经元掩码替代微调以增强鲁棒性。在前沿流匹配模型上的实验表明,该方法可有效擦除特定概念,同时保持模型整体性能,为更安全、负责任的生成模型提供可能。

原文摘要 · Abstract (English)

Recent advances in generative models have demonstrated remarkable capabilities in producing high-quality images, but their reliance on large-scale unlabeled data has raised significant safety and copyright concerns. Efforts to address these issues by erasing unwanted concepts have shown promise. However, many existing erasure methods involve excessive modifications that compromise the overall utility of the model. In this work, we address these issues by formulating a novel minimalist concept erasure objective based \emph{only} on the distributional distance of final generation outputs. Building on our formulation, we derive a tractable loss for differentiable optimization that leverages backpropagation through all generation steps in an end-to-end manner. We also conduct extensive analysis to show theoretical connections with other models and methods. To improve the robustness of the erasure, we incorporate neuron masking as an alternative to model fine-tuning. Empirical evaluations on state-of-the-art flow-matching models demonstrate that our method robustly erases concepts without degrading overall model performance, paving the way for safer and more responsible generative models.

生成模型概念擦除安全流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。