arXiv:2604.27414cs.CVcs.CR2026-04中稿 · SAE WCX 2026

研究视觉语言模型在自动驾驶中对抗攻击的跨架构迁移性,发现攻击效果高达91%。

Understanding Adversarial Transferability in Vision-Language Models for Autonomous Driving: A Cross-Architecture Analysis

论文配图:Understanding Adversarial Transferability in Vision-Language Models for Autonomous Driving: A Cross-Architecture Analysis
图 1 · 摘自论文原文
  • 通过真实路侧贴纸测试三种模型的对抗攻击迁移能力。
  • 跨架构攻击成功率73%-91%,关键决策窗口被干扰超64%。
  • 适用于关注自动驾驶安全性的研究人员和工程师。

视觉语言模型(VLMs)在自动驾驶中日益普及,因其结合了视觉感知与语言推理,支持更可解释的决策。然而,其对物理对抗攻击的鲁棒性,尤其是攻击是否能在不同VLM架构间迁移,尚不明确,这给实际应用带来风险——攻击者无需知晓车辆使用何种模型。本文通过系统性跨架构分析,评估了三种代表性架构(Dolphins、OmniDrive、LeapVAD),在人行横道和高速公路场景中,于路侧基础设施上部署可实现的物理贴纸进行测试。转移矩阵分析显示,跨架构攻击具有高度有效性:人行横道场景平均转移率TR=0.815,高速公路为0.833,最高达91%;即使贴纸未针对目标模型优化,仍能持续干扰64.7%-79.4%的关键决策窗口。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are increasingly used in autonomous driving because they combine visual perception with language-based reasoning, supporting more interpretable decision-making, yet their robustness to physical adversarial attacks, especially whether such attacks transfer across different VLM architectures, is not well understood and poses a practical risk when attackers do not know which model a vehicle uses. We address this gap with a systematic cross-architecture study of adversarial transferability in VLM-based driving, evaluating three representative architectures (Dolphins, OmniDrive, and LeapVAD) using physically realizable patches placed on roadside infrastructure in both crosswalk and highway scenarios. Our transfer-matrix evaluation shows high cross-architecture effectiveness, with transfer rates of 73-91% (mean TR = 0.815 for crosswalk and 0.833 for highway) and sustained frame-level manipulation over 64.7-79.4% of the critical decision window even when patches are not optimized for the target model.

自动驾驶对抗攻击视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。