让自动驾驶规划更懂时间,提升推理一致性与安全
From Prompts to Pavement Through Time: Temporal Grounding in Agentic Scene-to-Plan Reasoning

- 在多智能体通信中引入时间条件,增强动作连贯性
- 定量指标未显著提升,但质性分析显示预测危险与纠错能力增强
- 为时间感知的场景到规划推理提供首个实证基准
当前基于大语言模型(LLM)和大视觉语言模型(LMM)的自动驾驶高阶场景理解与规划方法仍忽视时间维度,导致连续动作推理不一致,影响安全与可解释性。本文探究在跨智能体通信中引入时间条件是否能保持或提升推理连贯性而不降低语义或逻辑一致性。通过设计三种逐步增强时间整合的规划器架构,在BDD-X数据集的精选子集上进行评估,采用语义、句法和逻辑指标。结果表明,尽管时间条件改变了推理风格,但对标准NLP正确性指标无统计显著提升。然而定性分析揭示了哨兵(Sentinel)模型具备预测性风险推理、稳定纠错行为及策略性分化能力。研究澄清了提示式时间锚定的局限性,并建立了首个面向时间感知的场景到规划推理的实证基准。
原文摘要 · Abstract (English)
Recent attempts to support high-level scene interpretation and planning in Autonomous Vehicles (AVs) using ensembles of Large Language Models (LLMs) and Large Multimodal Models (LMMs) continue to treat time as a secondary property. This lack of temporal grounding leads to inconsistencies in reasoning about continuous actions, undermining both safety and interpretability. This work explores whether temporal conditioning within inter-agent communication can preserve or enhance coherence without introducing degradation in semantic or logical consistency. To investigate this, we introduce three planner architectures with progressively increasing temporal integration and evaluate them on curated subsets of the BDD-X dataset using semantic, syntactic, and logical metrics. Results show that while temporal conditioning reshapes reasoning style, it yields no statistically significant improvements in standard NLP-based correctness metrics. However, qualitative analysis reveals predictive hazard reasoning, stable corrective behavior, and strategic divergence in the Sentinel. These findings clarify the limits of prompt-based temporal grounding and establish the first empirical benchmark for temporal scene-to-plan reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。