arXiv:2507.13598cs.CRcs.AI2025-07被引 2

GIFT让扩散模型免疫恶意微调,还能保留生成安全内容的能力。

GIFT: Gradient-aware Immunization of diffusion models against malicious Fine-Tuning with safe concepts retention

  • 通过梯度感知的双层优化,噪声化有害概念表示并保护安全数据表现。
  • 在恶意微调后,模型重学有害概念的能力显著下降,安全生成性能保持稳定。
  • 适合关注生成模型安全防御的研究者与开发者使用。

我们提出 GIFT:一种梯度感知的免疫技术,用于防御扩散模型在恶意微调下的攻击,同时保留其生成安全内容的能力。现有安全机制如安全检测器易被绕过,概念擦除方法在对抗性微调下失效。GIFT 将免疫过程建模为双层优化问题:上层目标通过表示噪声和最大化,削弱模型对有害概念的表征能力;下层目标则保持对安全数据的生成性能。实验表明,该方法能显著抑制模型在恶意微调后重新学习有害概念的能力,同时维持安全生成质量,为构建抗对抗性微调攻击的内在安全生成模型提供了新方向。

原文摘要 · Abstract (English)

We present GIFT: a {G}radient-aware {I}mmunization technique to defend diffusion models against malicious {F}ine-{T}uning while preserving their ability to generate safe content. Existing safety mechanisms like safety checkers are easily bypassed, and concept erasure methods fail under adversarial fine-tuning. GIFT addresses this by framing immunization as a bi-level optimization problem: the upper-level objective degrades the model's ability to represent harmful concepts using representation noising and maximization, while the lower-level objective preserves performance on safe data. GIFT achieves robust resistance to malicious fine-tuning while maintaining safe generative quality. Experimental results show that our method significantly impairs the model's ability to re-learn harmful concepts while maintaining performance on safe content, offering a promising direction for creating inherently safer generative models resistant to adversarial fine-tuning attacks.

扩散模型安全防御对抗性微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。