首次揭示文本生成图像模型中毒崩溃的内在机制
Understanding Implosion in Text-to-Image Generative Models
- 将跨注意力训练建模为图对齐问题,用对齐难度量化数据影响
- 对齐难度随中毒概念增多而上升,导致模型输出随机混乱图像
- 为扩散模型防御研究提供可解释分析工具,适合安全与可信AI研究者
近期研究发现,文本到图像生成模型极易受到多种投毒攻击。实证结果表明,仅改变个别文本提示与视觉特征的关联即可破坏模型。更严重的是,多重并发投毒攻击会引发“模型坍缩”——模型无法为未中毒提示生成有意义图像。这一现象凸显了缺乏直观理解此类攻击的理论框架。本文首次建立生成模型鲁棒性分析框架,通过建模和分析潜空间扩散模型中的交叉注意力机制,将交叉注意力训练视为一种抽象的“监督图对齐”问题,并正式定义对齐难度(AD)作为衡量训练数据影响的指标。AD越高,对齐越困难。我们证明:被污染的概念数量越多,AD越大;随着对齐任务加剧,模型输出逐渐失真,常将有意义文本映射到无意义或未定义的视觉表示,最终导致生成模型坍缩,大规模输出随机、不连贯图像。通过大量实验验证该分析框架,不仅解释了此前未解的模型坍缩现象,还揭示了新的深层洞察。本工作为研究扩散模型的投毒攻击及其防御提供了有力工具。
原文摘要 · Abstract (English)
Recent works show that text-to-image generative models are surprisingly vulnerable to a variety of poisoning attacks. Empirical results find that these models can be corrupted by altering associations between individual text prompts and associated visual features. Furthermore, a number of concurrent poisoning attacks can induce "model implosion," where the model becomes unable to produce meaningful images for unpoisoned prompts. These intriguing findings highlight the absence of an intuitive framework to understand poisoning attacks on these models. In this work, we establish the first analytical framework on robustness of image generative models to poisoning attacks, by modeling and analyzing the behavior of the cross-attention mechanism in latent diffusion models. We model cross-attention training as an abstract problem of "supervised graph alignment" and formally quantify the impact of training data by the hardness of alignment, measured by an Alignment Difficulty (AD) metric. The higher the AD, the harder the alignment. We prove that AD increases with the number of individual prompts (or concepts) poisoned. As AD grows, the alignment task becomes increasingly difficult, yielding highly distorted outcomes that frequently map meaningful text prompts to undefined or meaningless visual representations. As a result, the generative model implodes and outputs random, incoherent images at large. We validate our analytical framework through extensive experiments, and we confirm and explain the unexpected (and unexplained) effect of model implosion while producing new, unforeseen insights. Our work provides a useful tool for studying poisoning attacks against diffusion models and their defenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。