arXiv:2504.21307cs.CV2025-04被引 3

发现扩散模型删不掉的有害概念仍以可解释线性子空间形式存在,可攻可防。

The Linear Geometry of Interpretable Tokens: Jailbreaking Attacks and Defenses for Unlearned Diffusion Models

  • 通过学习正交攻击令牌,从文本嵌入空间读出残留有害概念的线性结构
  • 新攻击在多种条件下更有效且可迁移,防御机制能轻量抑制残留概念
  • 适合关注模型安全与可控生成的研究者,尤其关注去记忆技术漏洞者

扩散模型虽能生成高质量图像,但在提示下可能重现有害内容。尽管已有微调方法用于去记忆特定概念,但难以完全消除且保持其他概念生成质量,使模型易受越狱攻击。现有攻击方法揭示了这一漏洞,但对残留机制理解有限。本文发现,被删除的概念仍以可解释的线性子空间形式存在于令牌嵌入空间中,攻击与防御均可由此直接构建。我们提出SubAttack:通过学习一组正交攻击令牌嵌入(每个为人类可解读文本元素的线性组合),读出该子空间,证明未学模型仍通过相关文本成分保留目标概念。该攻击更具威力且在不同提示、初始噪声和模型间具有良好迁移性。反向地,投影移除该子空间即得SubDefense——一种轻量级即插即用防御机制,能有效抑制残留概念,同时比现有防御更强鲁棒性并更好保留安全生成质量。大量实验验证了其在多种去记忆方法、概念和攻击类型下的有效性,推动了对扩散模型去记忆漏洞的理解与缓解。

原文摘要 · Abstract (English)

Diffusion models excel at generating high-quality images but can memorize and reproduce harmful concepts when prompted. Although fine-tuning methods have been proposed to unlearn a target concept, they struggle to fully erase it while maintaining generation quality on other concepts, leaving models vulnerable to jailbreak attacks. Existing jailbreak methods demonstrate this vulnerability but offer limited insight into how unlearned models retain harmful concepts, limiting progress on effective defenses. In this work, we show that the erased concept persists as a coherent, interpretable linear subspace of the token embedding space, and that both an attack and a defense follow directly from this structure. We introduce SubAttack, a novel jailbreaking attack that reads out this subspace by learning an orthogonal set of attack token embeddings, each being a linear combination of human-interpretable textual elements, revealing that unlearned models still retain the target concept through related textual components. Furthermore, our attack is also more powerful and transferable across text prompts, initial noises, and unlearned models than prior attacks. Conversely, projecting out the same subspace yields SubDefense, a lightweight plug-and-play defense mechanism that suppresses the residual concept in unlearned models. SubDefense provides stronger robustness than existing defenses while better preserving safe generation quality. Extensive experiments across multiple unlearning methods, concepts, and attack types demonstrate that our approach advances both understanding and mitigation of vulnerabilities in diffusion unlearning.

扩散模型去记忆安全防御文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。