arXiv:2604.12833cs.CV2026-04

用可部署的光影干扰攻击视觉语言模型,让其产生语义错误。

Challenging Vision-Language Models with Physically Deployable Multimodal Semantic Lighting Attacks

  • 通过可控光照制造物理可实现的对抗性扰动
  • 使主流CLIP模型零样本分类性能下降,引发严重语义幻觉
  • 首次揭示VLM在真实场景下的语义级脆弱性,适合安全研究者参考

视觉语言模型(VLMs)表现出色,但其安全性尚未充分理解。现有对抗性研究几乎仅限于数字环境,忽视了真实世界威胁。随着VLMs在实际场景中广泛应用,对抗扰动必须具备物理可实现性,这一差距变得尤为关键。尽管具有现实意义,针对VLMs的物理攻击尚未系统研究。此类攻击可能导致识别失败并破坏多模态推理,引发下游任务中的严重语义误判。为此,我们提出首个可物理部署的对抗攻击框架——多模态语义光照攻击(MSLA)。MSLA利用可控对抗光照,在真实场景中干扰多模态语义理解,攻击语义对齐而非仅特定任务输出。结果表明,该方法显著降低主流CLIP变体的零样本分类性能,并在先进VLM如LLaVA和BLIP上引发严重的图像描述与视觉问答(VQA)语义幻觉。数字与物理域的大量实验验证了MSLA的有效性、迁移性与实用性。我们的发现首次证明VLMs极易受可部署语义攻击,暴露此前被忽略的鲁棒性缺口,凸显亟需开展VLMs在真实世界中的鲁棒性评估。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have shown remarkable performance, yet their security remains insufficiently understood. Existing adversarial studies focus almost exclusively on the digital setting, leaving physical-world threats largely unexplored. As VLMs are increasingly deployed in real environments, this gap becomes critical, since adversarial perturbations must be physically realizable. Despite this practical relevance, physical attacks against VLMs have not been systematically studied. Such attacks may induce recognition failures and further disrupt multimodal reasoning, leading to severe semantic misinterpretation in downstream tasks. Therefore, investigating physical attacks on VLMs is essential for assessing their real-world security risks. To address this gap, we propose Multimodal Semantic Lighting Attacks (MSLA), the first physically deployable adversarial attack framework against VLMs. MSLA uses controllable adversarial lighting to disrupt multimodal semantic understanding in real scenes, attacking semantic alignment rather than only task-specific outputs. Consequently, it degrades zero-shot classification performance of mainstream CLIP variants while inducing severe semantic hallucinations in advanced VLMs such as LLaVA and BLIP across image captioning and visual question answering (VQA). Extensive experiments in both digital and physical domains demonstrate that MSLA is effective, transferable, and practically realizable. Our findings provide the first evidence that VLMs are highly vulnerable to physically deployable semantic attacks, exposing a previously overlooked robustness gap and underscoring the urgent need for physical-world robustness evaluation of VLMs.

视觉语言模型对抗攻击物理安全语义攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。