arXiv:2507.07139cs.CVcs.CR2025-07中稿 · ICLR被引 6

用图像引导攻击生成模型遗忘机制,暴露其安全漏洞。

Image Can Bring Your Memory Back: A Novel Multi-Modal Guided Attack against Image Generation Model Unlearning

  • 通过参考图像优化对抗性图像提示,利用多模态特性发起攻击。
  • 在10种主流遗忘方法上均显著提升攻击成功率,保持语义一致性。
  • 揭示现有遗忘技术脆弱性,适合研究生成模型安全的学者参考。

近期基于扩散模型的图像生成模型(IGMs),如Stable Diffusion(SD),在生成视觉内容的质量与多样性方面取得显著进展。然而,其生成能力也引发了伦理、法律和社会层面的担忧,可能产生有害、误导或侵犯版权的内容。为缓解此类问题,机器遗忘(MU)作为一种有前景的解决方案,可选择性地从预训练模型中移除不良概念。然而,现有遗忘技术的鲁棒性与有效性仍缺乏充分探索,尤其在多模态对抗输入下。为此,我们提出Recall——一种专为破坏已遗忘IGMs鲁棒性设计的新型对抗框架。不同于依赖文本提示的现有方法,Recall利用扩散模型内在的多模态条件能力,通过单个语义相关的参考图像高效优化对抗性图像提示。在十种先进遗忘方法及多种任务上的广泛实验表明,Recall在对抗有效性、计算效率和原始文本提示语义保真度方面均显著优于现有基线。这些发现揭示了当前遗忘机制的关键漏洞,强调了构建更鲁棒解决方案以保障生成模型安全与可靠性的必要性。代码与数据公开于:https://github.com/ryliu68/RECALL。

原文摘要 · Abstract (English)

Recent advances in image generation models (IGMs), particularly diffusion-based architectures such as Stable Diffusion (SD), have markedly enhanced the quality and diversity of AI-generated visual content. However, their generative capability has also raised significant ethical, legal, and societal concerns, including the potential to produce harmful, misleading, or copyright-infringing content. To mitigate these concerns, machine unlearning (MU) emerges as a promising solution by selectively removing undesirable concepts from pretrained models. Nevertheless, the robustness and effectiveness of existing unlearning techniques remain largely unexplored, particularly in the presence of multi-modal adversarial inputs. To bridge this gap, we propose Recall, a novel adversarial framework explicitly designed to compromise the robustness of unlearned IGMs. Unlike existing approaches that predominantly rely on adversarial text prompts, Recall exploits the intrinsic multi-modal conditioning capabilities of diffusion models by efficiently optimizing adversarial image prompts with guidance from a single semantically relevant reference image. Extensive experiments across ten state-of-the-art unlearning methods and diverse tasks show that Recall consistently outperforms existing baselines in terms of adversarial effectiveness, computational efficiency, and semantic fidelity with the original textual prompt. These findings reveal critical vulnerabilities in current unlearning mechanisms and underscore the need for more robust solutions to ensure the safety and reliability of generative models. Code and data are publicly available at \textcolor{blue}{https://github.com/ryliu68/RECALL}.

生成模型对抗攻击机器遗忘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。