arXiv:2607.16938cs.CVcs.AI2026-07

用AI修复图像剔除物体,看清自动驾驶模型的决策依据。

What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning

论文配图:What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning
图 1 · 摘自论文原文
  • 通过生成填补技术移除图像中的物体,构造反事实场景。
  • 发现模型对行人、车辆影响大,对交通灯反应过度超出其视觉面积。
  • 揭示模型可能依赖人类无法理解的内部特征,适合安全评估与可信驾驶研究。

端到端自动驾驶模型能有效应对复杂交通场景,但其决策逻辑仍不透明。本文提出反事实消融框架CVAA,利用高保真生成修补技术从前视图像中系统性移除检测到的物体,构建反事实数据集以评估模型响应差异,从而分离各物体对规划行为的因果影响。在210个nuScenes场景上的实验表明,位于模型路径内的车辆和行人具有主导性影响,而交通灯虽视觉占比小,却产生超比例影响。但模型也对人类认为无关的物体做出强烈反应。为进一步探究,我们使用机制可解释性方法分析原始与修补图像在各模型层的中间表征变化。该双阶段方法实现从行为审计到表征理解的跨越,推动可解释自动驾驶系统发展,增强人机信任。

原文摘要 · Abstract (English)

End-to-end autonomous driving models are now able to navigate complex road scenarios, mapping raw sensor observations directly to observed paths for open-loop evaluation and often effective driving in closed-loop evaluation. Yet the internal logic of these safety-critical systems remains largely opaque, due to the complexity of traffic scenes. We propose a counterfactual ablation framework called Counterfactual Vision Action Analysis (CVAA) that systematically removes individual detected objects from front-camera images using photorealistic generative inpainting to prepare counterfactual sets to evaluate the difference in the model's response. This isolates the causal effect of each object's presence on the model's planning behaviour. Applied to the Alpamayo 1 trajectory predictor across 210 nuScenes driving scenes, we create a dataset Counter -nuScenes, using which we see that vehicles and pedestrians within the model's 'path' dominate causal influence as expected, while traffic lights, as expected, exert disproportionate effect relative to their image footprint. However, we also find cases where the model responds strongly to objects a human driver would consider irrelevant. This brings forth a deeper question: does the model itself view the scene as a sum of individual objects influencing the outcome, or does it encode an entirely different set of internal features that do not correspond to human-legible scene elements? To further understand this, we compare intermediate representations of original and inpainted image pairs using mechanistic interpretability techniques and examine the effect of the removal through the various model layers. Together, these two stages offer a path from behavioral auditing to representational understanding, creating explainable driving systems and solidifying human-AI trust.

自动驾驶可解释性视觉-语言-动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。