提出新方法让自动驾驶模型先推理再验证轨迹,减少错误联想。
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

- 用多选题形式预设轨迹候选,推理后才暴露真实路径
- 实验表明该方法显著降低幻觉,提升因果推理准确性
- 适合研究可验证自动驾驶决策的学者和工程师
当前自动驾驶视觉语言动作模型常采用思维链监督来增强推理能力,但现有标注流程普遍在教师模型前暴露未来轨迹真值。我们实证发现这导致轨迹锚定偏差:教师模型仅基于已知结果进行合理化,而非从场景证据推断决策,造成因果不一致的思维链和严重幻觉,尤其在因果复杂场景中更明显。移除真值轨迹可消除这一捷径,但开放生成轨迹会将高层决策与精确几何合成及低层动态耦合。为此,我们提出自动驾驶多选题(AD-MCQ),将规划转化为在显式轨迹候选中选择。进一步提出延迟暴露未来轨迹(DEFT-RLVR)方法,将未来轨迹从决策前锚点转为决策后验证目标。实验显示,DEFT-RLVR在保持甚至增强通用视觉能力的同时提升了自动驾驶推理性能。采用纯视觉语言模型推理,并可通过候选构造控制难度,AD-MCQ为可验证自动驾驶推理研究提供了灵活、可扩展的基础。
原文摘要 · Abstract (English)
Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision-Language Model (VLM) components, yet existing annotation pipelines commonly expose the teacher model to the logged ground-truth (GT) future trajectory. We empirically show that this induces trajectory anchoring bias: teacher models rationalize the revealed outcome rather than infer a decision from scene evidence, producing less causally faithful CoTs and substantially more severe hallucinations, especially in causally challenging scenes. Removing the GT trajectory eliminates this shortcut, but open-ended trajectory generation entangles high-level decision-making with precise geometric synthesis and low-level dynamics. To make trajectory-level driving decisions verifiable without requiring open-ended trajectory synthesis, we introduce Autonomous-Driving Multiple-Choice Question (AD-MCQ), which casts planning as selection among explicit trajectory candidates. Taking this a step further, we propose Deferred Exposure of Future Trajectories for RLVR (DEFT-RLVR) to transform future trajectories from pre-decision anchors into post-decision verification targets. Experimental results show that DEFT-RLVR improves AD reasoning while preserving or even enhancing general visual capabilities. With VLM-only inference and controllable difficulty through candidate construction, AD-MCQ provides a flexible, scalable, and extensible foundation for future research on verifiable AD reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。