arXiv:2605.31041cs.CVcs.AI2026-05

探究视觉信息如何影响视觉-语言-动作模型的驾驶行为。

Does Visual Information Play a Decisive Role in Vision-Language-Action Model Driving Behavior?

论文配图:Does Visual Information Play a Decisive Role in Vision-Language-Action Model Driving Behavior?
图 1 · 摘自论文原文
  • 设计多层级视觉扰动框架,分通道、信息、结构三方面测试视觉依赖。
  • 发现不同评估场景下模型对视觉的依赖程度差异显著。
  • 适合关注自动驾驶模型可解释性与安全性的研究者阅读。

视觉-语言-动作(VLA)模型在自动驾驶中展现出巨大潜力,表明统一多模态架构在联合建模感知与规划方面的前景。然而,当前VLA驱动行为如何依赖视觉信息仍不清晰。现有评估主要关注整体性能指标,缺乏系统化的诊断工具来量化视觉-行为依赖关系。本文提出一种结构化的多层级视觉扰动框架,系统分析VLA驱动模型中的视觉-行为依赖。该框架从通道级退化、信息级干扰和结构级修改三个互补维度施加受控视觉扰动,并应用于VLA驱动系统,在开环轨迹预测与闭环交互式安全评估中评估行为响应。实验结果揭示了评估依赖型的依赖模式,以及在抽象层次间不均衡的视觉奠基现象。这些发现呼吁更结构化的分析方法与原则化的设计,以更好理解视觉信息如何塑造行为,从而构建更安全、鲁棒的系统。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have demonstrated promising capability in autonomous driving, highlighting the potential of unified multimodal architectures for jointly modeling perception and planning. However, how current VLA-based driving behavior is grounded in visual information remains poorly understood. Existing evaluation protocols mainly focus on aggregate performance metrics, lacking structured and practical diagnostics to quantify visual-behavior dependency. In this work, we introduce a structured multi-level visual perturbation framework to analyze visual-behavior dependency in VLA-based driving models systematically. The framework organizes controlled visual perturbations along three complementary dimensions: channellevel degradation, information-level disruption, and structurelevel modification. We apply it to VLA-based driving systems and evaluate behavioral responses under both open-loop trajectory prediction and interactive closed-loop safety evaluation. Experimental results reveal evaluation-dependent dependency patterns and uneven visual grounding across abstraction levels. These findings call for more structured analyses and principled design of VLA driving models to better understand how visual information shapes behavior and develop safer, more robust systems.

自动驾驶多模态可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。