提出无需文本的攻击方法,发现去标识模型仍存视觉记忆。
TINA: Text-Free Inversion Attack for Unlearned Text-to-Image Diffusion Models
- 用无文本条件的反演攻击,绕过依赖文本的防御机制。
- 在主流去标识模型上成功复现被删除的概念图像。
- 揭示现有方法仅隐藏而非消除视觉知识,适合安全研究者参考。
尽管文本到图像扩散模型具有强大的生成能力,但概念擦除技术对于其安全部署至关重要,以防止生成有害内容。这催生了擦除防御与对抗性探测之间的动态博弈,推动了擦除方法的持续优化。然而,这种对抗性演进逐渐聚焦于以文本为中心的范式,即认为擦除就是切断文本到图像的映射,忽视了相关视觉知识在模型中依然存在。为验证这一观点,我们从视觉角度出发,利用DDIM反演探测被擦除概念是否仍有生成路径。然而,标准文本引导的DDIM反演会受到文本中心防御机制的抵制。为此,我们提出TINA(Text-free INversion Attack),通过在零文本条件下运行,规避现有文本中心防御。同时,TINA引入优化过程,克服因缺乏文本引导而累积的近似误差。实验表明,TINA可在采用先进去标识方法的模型中重新生成被擦除的概念。TINA的成功证明当前方法仅是掩盖概念,凸显了直接作用于内部视觉知识的新范式的紧迫性。
原文摘要 · Abstract (English)
Although text-to-image diffusion models exhibit remarkable generative power, concept erasure techniques are essential for their safe deployment to prevent the creation of harmful content. This has fostered a dynamic interplay between the development of erasure defenses and the adversarial probes designed to bypass them, and this co-evolution has progressively enhanced the efficacy of erasure methods. However, this adversarial co-evolution has converged on a narrow, text-centric paradigm that equates erasure with severing the text-to-image mapping, ignoring that the underlying visual knowledge related to undesired concepts still persist. To substantiate this claim, we investigate from a visual perspective, leveraging DDIM inversion to probe whether a generative pathway for the erased concept can still be found. However, identifying such a visual generative pathway is challenging because standard text-guided DDIM inversion is actively resisted by text-centric defenses within the erased model. To address this, we introduce TINA, a novel Text-free INversion Attack, which enforces this visual-only probe by operating under a null-text condition, thereby avoiding existing text-centric defenses. Moreover, TINA integrates an optimization procedure to overcome the accumulating approximation errors that arise when standard inversion operates without its usual textual guidance. Our experiments demonstrate that TINA regenerates erased concepts from models treated with state-of-the-art unlearning. The success of TINA proves that current methods merely obscure concepts, highlighting an urgent need for paradigms that operate directly on internal visual knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。