提出新攻击方法,揭露扩散模型去记忆的虚假安全
The Illusion of Forgetting: Attack Unlearned Diffusion via Initial Latent Variable Optimization
- 通过优化初始隐变量,重构被破坏的语言-知识映射
- 11种去记忆方法中,多数仍残留可被唤醒的隐藏知识
- 适合研究模型安全与去记忆机制的学者参考
文本到图像扩散模型常被用于生成有害或侵权内容,威胁公共利益。概念擦除(去记忆)是缓解此问题的有前景方案,但存在未明原因的‘遗忘幻觉’现象。基于实证分析,我们首次形式化解释其成因:多数去记忆仅部分破坏语言符号与内部知识间的映射关系,使知识以休眠记忆形式保留。我们进一步证明,去噪过程中的分布差异可作为映射保留程度的可测量指标,反映去记忆强度。受此启发,我们提出IVO(初始隐变量优化)攻击框架,通过优化初始隐变量,使去记忆模型的噪声分布与原始模型对齐,从而重建断裂的映射,唤醒休眠记忆。在11种去记忆技术及3个概念场景上的实验表明,IVO显著优于现有基线,暴露出当前去记忆机制的根本缺陷。警告:本文包含可能冒犯读者的不安全图像。
原文摘要 · Abstract (English)
Text-to-image diffusion models (DMs) are frequently abused to produce harmful or copyrighted content, violating public interests. Concept erasure (unlearning) is a promising paradigm to alleviate this issue. However, there exists a peculiar forgetting illusion phenomenon with unclear cause. Based on empirical analysis, we formally explain this cause: most unlearning partially disrupt the mapping between linguistic symbols and the underlying internal knowledge, leaving the knowledge intact as dormant memories. We further demonstrate that distributional discrepancy in the denoising process serves as a measurable indicator of how much of the mapping is retained, also reflecting unlearning strength. Inspired by this, we propose IVO (Initial Latent Variable Optimization), a novel attack framework designed to assess the robustness of current unlearning methods. IVO optimizes initial latent variables to realign the noise distribution of unlearned models with that of their vanilla counterparts, which reconstructs the fractured mappings and consequently revives dormant memories. Extensive experiments covering 11 unlearning techniques and 3 concept scenarios show that IVO outperforms state-of-the-art baselines, exposing fundamental flaws in current unlearning mechanisms. Warning: This paper has unsafe images that may offend some readers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。