arXiv:2509.20196cs.CVcs.LG2025-09被引 4

提出首个通用伪装攻击框架,让自动驾驶视觉语言模型误判路况指令。

Universal Camouflage Attack on Vision-Language Models for Autonomous Driving

  • 在特征空间生成可物理实现的伪装纹理,跨模型通用性强。
  • 在3类场景下攻击成功率提升30%,显著超越现有方法。
  • 适用于真实道路视角变化与动态环境,适合安全测试场景。

面向自动驾驶的视觉语言建模(VLM-AD)虽具备强大的多模态推理能力,但易受对抗攻击威胁。现有攻击存在两大缺陷:物理攻击多针对视觉模块,难以迁移至多模态系统;数字攻击则集中于特定指令,泛化性差。为此,本文提出首个通用伪装攻击(UCA)框架,不依赖对齐层优化,而是在特征空间生成具有强泛化能力的物理可实现伪装纹理。基于对编码器与投影层脆弱性的观察,引入特征分歧损失(FDL),最大化干净图像与对抗图像之间的表征差异。同时采用多尺度学习策略并动态调整采样率,增强对实际场景中尺度与视角变化的适应性,提升训练稳定性。大量实验表明,UCA可在多种VLM-AD模型与驾驶场景中引发错误驾驶指令,3-P指标性能较现有最优方法提升30%以上。且在不同视角与动态条件下仍具强鲁棒性,具备实际部署潜力。

原文摘要 · Abstract (English)

Visual language modeling for automated driving is emerging as a promising research direction with substantial improvements in multimodal reasoning capabilities. Despite its advanced reasoning abilities, VLM-AD remains vulnerable to serious security threats from adversarial attacks, which involve misleading model decisions through carefully crafted perturbations. Existing attacks have obvious challenges: 1) Physical adversarial attacks primarily target vision modules. They are difficult to directly transfer to VLM-AD systems because they typically attack low-level perceptual components. 2) Adversarial attacks against VLM-AD have largely concentrated on the digital level. To address these challenges, we propose the first Universal Camouflage Attack (UCA) framework for VLM-AD. Unlike previous methods that focus on optimizing the logit layer, UCA operates in the feature space to generate physically realizable camouflage textures that exhibit strong generalization across different user commands and model architectures. Motivated by the observed vulnerability of encoder and projection layers in VLM-AD, UCA introduces a feature divergence loss (FDL) that maximizes the representational discrepancy between clean and adversarial images. In addition, UCA incorporates a multi-scale learning strategy and adjusts the sampling ratio to enhance its adaptability to changes in scale and viewpoint diversity in real-world scenarios, thereby improving training stability. Extensive experiments demonstrate that UCA can induce incorrect driving commands across various VLM-AD models and driving scenarios, significantly surpassing existing state-of-the-art attack methods (improving 30\% in 3-P metrics). Furthermore, UCA exhibits strong attack robustness under diverse viewpoints and dynamic conditions, indicating high potential for practical deployment.

对抗攻击自动驾驶视觉语言模型安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。