让车与路侧协同推理,提升自动驾驶长距决策能力。
DH-VLM: Dual-Horizon Cooperative Latent Reasoning for Autonomous Driving

- 路侧与车辆分层协作,通过潜空间引导实现信息互补
- 规划误差降低14.6%,碰撞率下降26.9%,性能达新高
- 通信开销减少57.3%,适合实际部署的鲁棒协同系统
大规模语言模型可增强自动驾驶的全局理解与长时序规划能力,但孤立车辆受感知范围有限和遮挡影响,难以可靠决策;且计算与延迟成本过高,难以车载部署。协同驾驶通过外部智能体信息共享提供解决方案,但现有方法在实际约束下语义推理能力不足。为此,我们提出DH-VLM——一种双时域协同潜空间推理框架,实现基础设施与自车间的非对称语义协作。基础设施聚合多层隐状态,生成全局推理潜空间引导信号,并通过“基础设施驱动潜空间演化”机制注入自车模型,实现条件化潜变量优化。该设计使自车在保持本地规划自主性的同时,获得长距离上下文理解能力。此外,我们构建了面向协同的问答数据集(QA),覆盖基础场景理解与自车个性化认知,支持反事实与安全导向推理。大量实验表明,DH-VLM在规划任务上达到当前最优表现:相较前人最佳方法,L2误差降低14.6%,碰撞率下降26.9%。相比基于查询的端到端协同方法,通信成本减少57.3%,GPU内存使用量降低25.5%,同时具备对基础设施引导错误的强鲁棒性,为协同自动驾驶提供高效可靠的实用范式。
原文摘要 · Abstract (English)
Large-scale language models for autonomous driving enable enhanced global understanding and long-horizon planning. However, when deployed in isolated vehicles, limited sensing range and occlusions restrict reliable decision-making, and the substantial computational and latency overhead makes on-board deployment impractical. Cooperative driving provides a potential solution by leveraging external agents for information exchange, but existing methods remain limited in semantic reasoning capability under practical constraints. To address these challenges, we propose DH-VLM, a dual-horizon cooperative latent reasoning framework that enables asymmetric semantic cooperation between the infrastructure and ego vehicle. The infrastructure aggregates multi-layer hidden states to form a global-reasoning horizon latent guidance, which is integrated into the ego model through an Infrastructure-Driven Latent Evolution mechanism for conditional latent refinement. This enables the ego vehicle to leverage long-range contextual understanding while preserving autonomous decision-making within its local planning horizon. Furthermore, we construct a cooperation-oriented question-answer (QA) dataset covering fundamental scene understanding and ego-personalized comprehension to support counterfactual and safety-aware reasoning. Extensive experiments demonstrate that DH-VLM achieves state-of-the-art planning performance, outperforming the previous state of the art by 14.6% in L2 error and 26.9% in collision rate. Compared with query-based end-to-end cooperative driving methods, our approach reduces the communication cost by 57.3% and GPU memory usage by 25.5%, while maintaining strong robustness against infrastructure guidance errors, providing a practical and robust paradigm for cooperative autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。