arXiv:2608.21032cs.RO2026-08

构建支持语言交互的协同自动驾驶评估平台,实现高精度闭环推理。

Roadside-Cooperative Autonomous Driving: From Data Platform to Vision-Language End-to-End Reasoning

论文配图:Roadside-Cooperative Autonomous Driving: From Data Platform to Vision-Language End-to-End Reasoning
图 1 · 摘自论文原文
  • 设计双视角感知融合模块,解决车端与路侧视角差异问题。
  • 在遮挡场景下实现98.21%路线完成率和76.02分驾驶评分。
  • 适合研究协同感知、视觉语言模型与智能交通系统的学者。

车联网(V2X)协作可实现超视距感知,缓解单车感知中的遮挡问题。然而,现有V2X基准数据集缺乏闭环评估支持与语言引导监督,制约了视觉语言模型(VLM)在端到端协同驾驶中的发展。为此,我们提出V2XBench仿真平台,具备同步车端-路侧感知与闭环评估能力,并构建了逐步结构化的对话式VQA数据集Chat-V2XBench。基于此,我们提出AURORA端到端协同驾驶框架:采用双视角感知架构,通过查询级跨视图对齐与融合(CQAF)模块缓解车端与路侧视角间的空间与语义差异;利用统一特征表示,结合LoRA微调的VLM实现语义推理与生成式轨迹规划。在V2XBench上的大量闭环测试表明,AURORA在严重遮挡场景中达到98.21%的路线完成率与76.02分驾驶评分,且路侧通信带宽需求低。该工作开创了可扩展的V2X-VLM协同驾驶范式,为下一代智能网联自动驾驶提供新路径。

原文摘要 · Abstract (English)

Vehicle-to-Everything (V2X) cooperation enables beyond-line-of-sight perception, mitigating occlusions in single-vehicle sensing. However, existing V2X benchmarks provide limited support for closed-loop evaluation and language-grounded supervision, hindering the development of vision-language models (VLMs) for end-to-end cooperative driving. To address these limitations, we introduce V2XBench, a simulation platform featuring synchronized ego--roadside sensing and closed-loop evaluation, together with Chat-V2XBench, a progressively structured VQA dataset for cooperative reasoning. Building upon this benchmark infrastructure, we propose AURORA, an end-to-end cooperative driving framework. Equipped with a dual-view perception architecture, AURORA mitigates spatial and semantic discrepancies across ego and roadside viewpoints through a query-level Cross-View Query Alignment and Fusion (CQAF) module. Leveraging the resulting unified tokens, a LoRA-adapted VLM bridges semantic reasoning and generative trajectory planning. Extensive closed-loop evaluations on V2XBench demonstrate that AURORA achieves state-of-the-art performance in heavily occluded scenarios, with a Route Completion rate of 98.21% and a Driving Score of 76.02, while requiring low roadside communication bandwidth. Ultimately, this work pioneers an extensible V2X--VLM paradigm, paving the way for next-generation cooperative autonomous driving.

协同驾驶视觉语言模型V2X闭环评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。