让自动驾驶理解指令前先预演场景,提升定位准确性。
Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles
- 用世界模型预演未来空间状态,辅助理解指令
- 在六个基准上领先,长文本和多智能体场景表现突出
- 适合需要强推理的自动驾驶交互任务
将自然语言指令转化为目标物体定位对自动驾驶至关重要。现有方法因缺乏对三维空间关系和场景演化预期的推理,在模糊或上下文依赖的指令上表现不佳。受世界模型启发,我们提出ThinkDeeper框架,通过空间感知世界模型(SA-WM)将当前场景提炼为指令相关隐状态,并滚动生成一系列未来隐状态,提供前瞻线索以消除歧义。同时,超图引导解码器分层融合这些状态与多模态输入,捕捉高阶空间依赖关系,实现鲁棒定位。此外,我们构建了DrivePilot数据集,采用检索增强生成与思维链提示的LLM流水线生成语义标注。在六个基准上的实验证明,ThinkDeeper在Talk2Car榜单排名第一,优于现有基线模型,在DrivePilot、MoCAD及RefCOCO/+/g上均取得最佳性能。尤其在长文本、多智能体和模糊场景中表现出强鲁棒性,且仅用50%数据训练仍保持优异表现。
原文摘要 · Abstract (English)
Interpreting natural-language commands to localize target objects is critical for autonomous driving (AD). Existing visual grounding (VG) methods for autonomous vehicles (AVs) typically struggle with ambiguous, context-dependent instructions, as they lack reasoning over 3D spatial relations and anticipated scene evolution. Grounded in the principles of world models, we propose ThinkDeeper, a framework that reasons about future spatial states before making grounding decisions. At its core is a Spatial-Aware World Model (SA-WM) that learns to reason ahead by distilling the current scene into a command-aware latent state and rolling out a sequence of future latent states, providing forward-looking cues for disambiguation. Complementing this, a hypergraph-guided decoder then hierarchically fuses these states with the multimodal input, capturing higher-order spatial dependencies for robust localization. In addition, we present DrivePilot, a multi-source VG dataset in AD, featuring semantic annotations generated by a Retrieval-Augmented Generation (RAG) and Chain-of-Thought (CoT)-prompted LLM pipeline. Extensive evaluations on six benchmarks, ThinkDeeper ranks #1 on the Talk2Car leaderboard and surpasses state-of-the-art baselines on DrivePilot, MoCAD, and RefCOCO/+/g benchmarks. Notably, it shows strong robustness and efficiency in challenging scenes (long-text, multi-agent, ambiguity) and retains superior performance even when trained on 50% of the data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。