发现视觉语言驾驶模型的推理与规划之间存在因果断联。
More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models
- 构建驱动思维数据集DriveMind,分离先验与待推理信号。
- 移除先验导致规划性能大幅下降,但移除推理影响甚微。
- 提出无需训练的诊断工具,可检测模型对先验的依赖程度。
视觉语言模型(VLM)驾驶代理通过生成自然语言推理来实现可解释的端到端自主驾驶,但其规划是否由推理因果驱动仍未经验证。为此,我们构建了DriveMind——一个大规模驾驶视觉问答数据集,其计划对齐的思维链(CoT)由nuPlan自动生成。数据生成过程将传感器与标注转换为结构化输入,并关键性地分离先验与待推理信号,支持纯净的信息消融实验。基于DriveMind,我们使用监督微调(SFT)和组相对策略优化(GRPO)训练代表性VLM代理,并在nuPlan指标下评估。结果表明,推理-规划间存在持续的因果断联:移除自身/导航先验导致规划得分显著下降,而移除CoT仅带来微小变化。注意力分析进一步显示,规划主要关注先验而非CoT。据此,我们提出推理-规划解耦假说,认为训练所得推理仅为附属产物而非因果中介。为实现高效诊断,我们还引入一种新型无训练探针,通过评估模型在微小输入扰动下的规划鲁棒性,测量其对先验的依赖程度。综上,我们为社区提供新数据集与诊断工具,以评估未来模型的因果保真度。
原文摘要 · Abstract (English)
Vision-Language Model (VLM) driving agents promise explainable end-to-end autonomy by first producing natural-language reasoning and then predicting trajectory planning. However, whether planning is causally driven by this reasoning remains a critical but unverified assumption. To investigate this, we build DriveMind, a large-scale driving Visual Question Answering (VQA) corpus with plan-aligned Chain-of-Thought (CoT), automatically generated from nuPlan. Our data generation process converts sensors and annotations into structured inputs and, crucially, separates priors from to-be-reasoned signals, enabling clean information ablations. Using DriveMind, we train representative VLM agents with Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO) and evaluate them with nuPlan's metrics. Our results, unfortunately, indicate a consistent causal disconnect in reasoning-planning: removing ego/navigation priors causes large drops in planning scores, whereas removing CoT produces only minor changes. Attention analysis further shows that planning primarily focuses on priors rather than the CoT. Based on this evidence, we propose the Reasoning-Planning Decoupling Hypothesis, positing that the training-yielded reasoning is an ancillary byproduct rather than a causal mediator. To enable efficient diagnosis, we also introduce a novel, training-free probe that measures an agent's reliance on priors by evaluating its planning robustness against minor input perturbations. In summary, we provide the community with a new dataset and a diagnostic tool to evaluate the causal fidelity of future models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。