arXiv:2606.12706cs.CV2026-06

评测视觉语言模型的思考过程与驾驶动作的关系,发现思考未必真有用。

VLADriveBench: Evaluating CoT-Action Relationship in VLA for Autonomous Driving

  • 设计新评测框架,结合观察指标和干预实验评估思考过程与动作关联
  • 发现部分模型思考与驾驶无关,另一些模型思考直接影响驾驶决策
  • 适合关注自动驾驶模型可解释性与安全性的研究者使用

视觉-语言-动作(VLA)模型在生成驾驶轨迹的同时会输出链式思考(CoT),但现有基准仅评估轨迹质量,未检验思考是否相关、一致或具有因果性。本文提出VLADriveBench,融合观测指标(提及、幻觉、矛盾、动作对齐)与CoT干预协议,从多角度评估思考与动作的关系。在两个架构的三款模型上测试发现,两种分析结果可能截然不同:ORION在观测对齐上得分最高,但其思考为附带现象;Alpamayo v1.5得分较低,但其思考具强因果性,且视觉显著性调控了思考的影响程度。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models generate chain-of-thought (CoT) reasoning alongside driving trajectories, but existing benchmarks evaluate only trajectory quality and do not assess whether the CoT is relevant, consistent, or causally connected to the driving action. We introduce VLADriveBench, a framework that combines observational metrics (mentioning, hallucination, contradiction, action alignment) with a CoT intervention protocol to provide complementary views of the CoT-action relationship. Applying VLADriveBench to three models across two architectures, we find that the two analyses can diverge sharply: ORION scores highest on observational alignment yet its CoT is epiphenomenal, while Alpamayo v1.5 scores lower yet its CoT is strongly causal, with visual salience gating the extent of CoT influence.

自动驾驶模型可解释性思考过程评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。