提出新方法破解概念擦除漏洞,揭示模型嵌入空间的固有风险
Rethinking the Vulnerability of Concept Erasure and a New Method
- 基于坐标下降法设计新型恢复算法,突破现有攻击性能上限
- 实验显示新方法效果比现有方法强17.8倍,且在多种设置下稳定有效
- 适合关注生成模型安全性的研究人员,尤其关注隐私保护与对抗攻击
文本到图像扩散模型的广泛应用引发了版权和安全方面的重大担忧,尤其是在生成受版权保护或有害图像方面。为此,概念擦除(防御)方法被提出,通过后训练微调“遗忘”特定概念。然而,近期的概念恢复(攻击)方法表明,这些被擦除的概念可通过对抗性提示重新恢复,暴露出当前防御机制的关键漏洞。本文首先探究了对抗性漏洞的根本来源,发现漏洞普遍存在于概念擦除模型的提示嵌入空间中,这一特性源自原始预训练模型。此外,我们提出**RECORD**,一种基于坐标下降的新型恢复算法,在多项实验中表现优于现有方法高达17.8倍,并评估了其计算-性能权衡关系,提出了加速策略。
原文摘要 · Abstract (English)
The proliferation of text-to-image diffusion models has raised significant privacy and security concerns, particularly regarding the generation of copyrighted or harmful images. In response, concept erasure (defense) methods have been developed to "unlearn" specific concepts through post-hoc finetuning. However, recent concept restoration (attack) methods have demonstrated that these supposedly erased concepts can be recovered using adversarially crafted prompts, revealing a critical vulnerability in current defense mechanisms. In this work, we first investigate the fundamental sources of adversarial vulnerability and reveal that vulnerabilities are pervasive in the prompt embedding space of concept-erased models, a characteristic inherited from the original pre-unlearned model. Furthermore, we introduce **RECORD**, a novel coordinate-descent-based restoration algorithm that consistently outperforms existing restoration methods by up to 17.8 times. We conduct extensive experiments to assess its compute-performance tradeoff and propose acceleration strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。