提出AEGIS框架,实现扩散模型概念擦除的强鲁棒性与高保留性统一
AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion Models
- 基于对抗性梯度协同设计,无需训练数据即可实现概念擦除
- 在不牺牲无关概念保留的前提下,有效抵抗语义相关提示攻击
- 适合需安全可控生成的工业级应用,如内容审核与合规生成
概念擦除可阻止扩散模型生成有害内容;但现有方法面临鲁棒性与保留性的权衡。鲁棒性指经概念擦除微调后的模型在语义相关提示下仍能抵抗被擦除概念的重新激活;保留性指无关概念得以保持,以维持模型整体可用性。二者在实践中均至关重要,但同时优化极为困难,因现有工作常提升一方而牺牲另一方。例如,将单个被擦除提示映射到固定安全目标会留下类别级漏洞,易被提示攻击利用;而侧重保留的方案在面对自适应攻击时表现不佳。本文提出对抗性梯度协同擦除(AEGIS),一种无需训练数据的框架,在不降低保留性的同时显著提升鲁棒性。
原文摘要 · Abstract (English)
Concept erasure helps stop diffusion models (DMs) from generating harmful content; but current methods face robustness retention trade off. Robustness means the model fine-tuned by concept erasure methods resists reactivation of erased concepts, even under semantically related prompts. Retention means unrelated concepts are preserved so the model's overall utility stays intact. Both are critical for concept erasure in practice, yet addressing them simultaneously is challenging, as existing works typically improve one factor while sacrificing the other. Prior work typically strengthens one while degrading the other, e.g., mapping a single erased prompt to a fixed safe target leaves class level remnants exploitable by prompt attacks, whereas retention-oriented schemes underperform against adaptive adversaries. This paper introduces Adversarial Erasure with Gradient Informed Synergy (AEGIS), a retention-data-free framework that advances both robustness and retention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。