多模态生成模型的后门攻击常只依赖单一模态,其他模态形同虚设。
When One Modality Rules Them All: Backdoor Modality Collapse in Multimodal Diffusion Models
- 发现后门攻击会集中于少数模态,出现'赢家通吃'现象
- 跨模态交互几乎为零甚至有害,违背协同攻击直觉
- 适合关注多模态模型安全性的研究者阅读
尽管扩散模型已革新视觉内容生成,但其快速应用凸显了漏洞研究的紧迫性,如后门攻击。在多模态扩散模型中,人们普遍认为同时攻击文本与图像等多模态可增强整体后门效果。本文挑战此假设,揭示后门模态坍塌现象:后门机制退化为高度依赖少数模态,其余模态变得冗余。为此,我们提出两个新指标——触发模态归因(TMA)与跨触发交互(CTI),在多种训练配置下对多模态条件扩散模型进行广泛实验。结果一致显示后门行为呈现‘赢家通吃’动态。发现:(1) 攻击常退化为子集模态主导;(2) 跨模态交互微弱甚至负面,与协同脆弱性直觉相悖。这些结果揭示当前评估中的关键盲点——高攻击成功率常掩盖对少数模态的深层依赖。本研究为机制分析与未来防御开发提供了坚实基础。
原文摘要 · Abstract (English)
While diffusion models have revolutionized visual content generation, their rapid adoption has underscored the critical need to investigate vulnerabilities, e.g., to backdoor attacks. In multimodal diffusion models, it is natural to expect that attacking multiple modalities simultaneously (e.g., text and image) would yield complementary effects and strengthen the overall backdoor. In this paper, we challenge this assumption by investigating the phenomenon of Backdoor Modality Collapse, a scenario where the backdoor mechanism degenerates to rely predominantly on a subset of modalities, rendering others redundant. To rigorously quantify this behavior, we introduce two novel metrics: Trigger Modality Attribution (TMA) and Cross-Trigger Interaction (CTI). Through extensive experiments across diverse training configurations in multimodal conditional diffusion, we consistently observe a ``winner-takes-all'' dynamic in backdoor behavior. Our results reveal that (1) attacks often collapse into subset-modality dominance, and (2) cross-modal interaction is negligible or even negative, contradicting the intuition of synergistic vulnerability. These findings highlight a critical blind spot in current assessments, suggesting that high attack success rates often mask a fundamental reliance on a subset of modalities. This establishes a principled foundation for mechanistic analysis and future defense development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。