检验视觉语言模型在自动驾驶中的推理可靠性,发现其常因记忆模式而非真实时序理解导致答案不一致。
Probing the Reliability of Driving VLMs: From Inconsistent Responses to Grounded Temporal Reasoning
- 通过链式思考自监督微调提升模型时序推理一致性
- 新构建FutureVQA数据集评估未来场景推理能力
- 强视觉理解不等于好时序推理,模型易依赖预训练模式
可靠的驾驶助手应基于可观测信息进行时序一致的推理。本文研究视觉语言模型(VLMs)作为驾驶助手时是否能保持响应一致性,并理解当前观察如何影响未来判断,还是仅反映训练中记忆的模式而缺乏时序根基推理。尽管近期将VLMs融入自动驾驶的研究侧重场景理解与指令生成,普遍假设强视觉解析自然带来一致未来推理,从而确保可靠决策,但本文对此提出质疑。我们聚焦两大限制可靠性的问题:响应不一致——微小输入扰动引发不同回答,甚至退化为近似随机猜测;以及有限的时序推理能力——模型无法从当前观测中推断并对齐连续事件,常产生错误或矛盾回应。此外,我们发现具备强视觉理解的模型在需时序推理的任务上未必表现最佳,表明其倾向于过度依赖预训练模式而非建模时间动态。为解决此问题,我们采用现有评估方法,并引入FutureVQA——一个由人工标注的基准数据集,专门用于评估未来场景推理。同时,提出一种无需时序标签的简单有效自监督微调方法,结合思维链推理,显著提升一致性与时序推理能力。
原文摘要 · Abstract (English)
A reliable driving assistant should provide consistent responses based on temporally grounded reasoning derived from observed information. In this work, we investigate whether Vision-Language Models (VLMs), when applied as driving assistants, can response consistantly and understand how present observations shape future outcomes, or whether their outputs merely reflect patterns memorized during training without temporally grounded reasoning. While recent efforts have integrated VLMs into autonomous driving, prior studies typically emphasize scene understanding and instruction generation, implicitly assuming that strong visual interpretation naturally enables consistant future reasoning and thus ensures reliable decision-making, a claim we critically examine. We focus on two major challenges limiting VLM reliability in this setting: response inconsistency, where minor input perturbations yield different answers or, in some cases, responses degenerate toward near-random guessing, and limited temporal reasoning, in which models fail to reason and align sequential events from current observations, often resulting in incorrect or even contradictory responses. Moreover, we find that models with strong visual understanding do not necessarily perform best on tasks requiring temporal reasoning, indicating a tendency to over-rely on pretrained patterns rather than modeling temporal dynamics. To address these issues, we adopt existing evaluation methods and introduce FutureVQA, a human-annotated benchmark dataset specifically designed to assess future scene reasoning. In addition, we propose a simple yet effective self-supervised tuning approach with chain-of-thought reasoning that improves both consistency and temporal reasoning without requiring temporal labels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。