arXiv:2604.08395cs.CVcs.AI2026-04

提出可自适应上下文的视觉语言模型后门攻击,隐蔽性更强。

Phantasia: Context-Adaptive Backdoors in Vision Language Models

  • 根据输入语境动态生成恶意响应,而非固定模式。
  • 在多种防御下仍保持高成功率,且不影响正常性能。
  • 揭示现有攻击易被检测,推动安全研究进展。

视觉语言模型(VLMs)在多模态理解方面取得显著进展,但其安全性尤其是后门攻击漏洞仍严重缺乏研究。现有攻击大多依赖生成具有固定特征的中毒输出,容易被检测。本文首次证明,现有VLM后门攻击的隐蔽性被严重高估——通过迁移其他领域(如纯视觉或纯文本)的防御技术,可轻易识别多个前沿攻击。为弥补此缺陷,我们提出Phantasia:一种上下文自适应后门攻击,能动态调整中毒输出语义以匹配输入上下文,生成看似合理但恶意的响应,显著提升隐蔽性与适应性。在多种VLM架构上的实验表明,Phantasia在各类防御设置下均实现最高攻击成功率,同时保持正常性能。

原文摘要 · Abstract (English)

Recent advances in Vision-Language Models (VLMs) have greatly enhanced the integration of visual perception and linguistic reasoning, driving rapid progress in multimodal understanding. Despite these achievements, the security of VLMs, particularly their vulnerability to backdoor attacks, remains significantly underexplored. Existing backdoor attacks on VLMs are still in an early stage of development, with most current methods relying on generating poisoned responses that contain fixed, easily identifiable patterns. In this work, we make two key contributions. First, we demonstrate for the first time that the stealthiness of existing VLM backdoor attacks has been substantially overestimated. By adapting defense techniques originally designed for other domains (e.g., vision-only and text-only models), we show that several state-of-the-art attacks can be detected with surprising ease. Second, to address this gap, we introduce Phantasia, a context-adaptive backdoor attack that dynamically aligns its poisoned outputs with the semantics of each input. Instead of producing static poisoned patterns, Phantasia encourages models to generate contextually coherent yet malicious responses that remain plausible, thereby significantly improving stealth and adaptability. Extensive experiments across diverse VLM architectures reveal that Phantasia achieves state-of-the-art attack success rates while maintaining benign performance under various defensive settings.

后门攻击视觉语言模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。