提出新方法检测扩散模型中被删除概念的视觉残留知识。
TINA+: Probing Residual Visual Knowledge in Unlearned Diffusion Models via Diffusion-Consistent Text-Free Inversion

- 用无文本反演技术追踪生成路径,探测视觉知识是否残留
- 发现未约束反演可能产生虚假轨迹,误判残留知识
- 引入扩散一致性正则化,提升探测可靠性,适合安全评估
尽管文本到图像扩散模型生成能力强大,概念擦除技术对防止有害内容至关重要。现有对抗性探针主要依赖文本路径测试擦除效果,但忽视了视觉知识是否仍存在。本文通过扩散反演技术,从视觉角度探究被擦除概念能否重建。在无文本条件下,标准反演因避开文本路径而放大近似误差,难以忠实恢复生成轨迹。为此,提出TINA+,一种基于优化的无文本反演攻击,并引入扩散一致性轨迹正则化,抑制随机初始化模型也能重建目标概念的虚假轨迹。该正则化通过惩罚偏离预期能量演化路径的轨迹,有效消除伪迹。在十二种擦除方法、四个概念擦除任务和多种模型架构上的实验表明,TINA+能可靠探测残留视觉知识。结果表明,当前方法多通过切断文本-图像关联来隐藏概念,而非真正消除视觉知识。
原文摘要 · Abstract (English)
Although text-to-image diffusion models exhibit remarkable generative power, concept erasure techniques are essential for preventing harmful content. Existing adversarial probes evaluate these methods by testing whether erased concepts can still be recovered. However, existing erasure and probe methods remain largely text-centric, focusing on whether the text-to-image mapping is severed while overlooking whether the corresponding visual knowledge remains. To investigate this question from a visual perspective, we leverage diffusion inversion to probe whether a generative trajectory can reconstruct visual instances of an erased concept. Under a null-text condition, standard inversion avoids the textual pathway but amplifies approximation errors, hindering faithful trajectory recovery. To address this challenge, we introduce TINA+, a diffusion-consistent Text-free INversion Attack equipped with optimization-based inversion. We also find that unconstrained diffusion inversion may discover spurious trajectories, even allowing a randomly initialized diffusion model to reconstruct the target concept. Such trajectories may falsely indicate residual visual knowledge. TINA+ therefore introduces Diffusion-Consistent Trajectory Regularization to suppress this failure mode. By penalizing trajectories that fall far below the expected marginal energy evolution of diffusion, TINA+ suppresses spurious inversion paths while preserving its ability to recover erased concepts. Experiments across twelve erasure methods, four concept-erasure tasks, and different model architectures demonstrate that TINA+ reliably probes residual visual knowledge through diffusion-consistent visual trajectories. These results provide stronger evidence that current methods often obscure concepts by severing text-image links rather than eliminating the underlying visual knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。