arXiv:2604.02710cs.ROcs.AI2026-04被引 4

构建多视角自动驾驶评测数据集,推动模型跨车端、路侧与协同视角的智能推理研究。

V2X-QA: A Comprehensive Reasoning Dataset and Benchmark for Multimodal Large Language Models in Autonomous Driving Across Ego, Infrastructure, and Cooperative Views

  • 设计视点解耦评估框架,在统一问答体系中对比车辆、路侧与协同视角表现。
  • 实测显示路侧视点更利于宏观交通理解,协同推理需跨视角对齐而非简单叠加视觉信息。
  • 提出V2X-MoE模型,通过显式视点路由和专用专家模块提升多视角协作能力,适合自动驾驶研究者参考。

多模态大语言模型在自动驾驶领域展现出巨大潜力,但现有评测基准仍以车辆为中心,难以系统评估路侧与协同驾驶场景下的模型性能。本文提出V2X-QA,一个基于真实世界数据的多视角评测数据集与基准,涵盖车端、路侧及协同视角。该基准采用视点解耦的评估协议,在统一的多选问答(MCQA)框架下,支持车辆独有、路侧独有及协同驾驶条件下的可控对比。任务涵盖感知、预测与推理规划共十二类,经专家验证标注,实现对视点依赖能力的细粒度诊断。在十种主流开源与专有模型上的测试表明,视点可访问性显著影响性能,路侧视角有助于宏观交通理解;而协同推理仍具挑战,因其需要跨视角对齐与证据融合,非仅增加视觉输入即可解决。为此,我们提出与基准对齐的V2X-MoE基线模型,采用显式视点路由与视角专用LoRA专家。其优异表现表明,显式视点专业化是多视角推理的可行方向。V2X-QA为研究连通自动驾驶中的多视角推理、可靠性与协同物理智能提供了基础。数据集与V2X-MoE资源已公开:https://github.com/junwei0001/V2X-QA。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have shown strong potential for autonomous driving, yet existing benchmarks remain largely ego-centric and therefore cannot systematically assess model performance in infrastructure-centric and cooperative driving conditions. In this work, we introduce V2X-QA, a real-world dataset and benchmark for evaluating MLLMs across vehicle-side, infrastructure-side, and cooperative viewpoints. V2X-QA is built around a view-decoupled evaluation protocol that enables controlled comparison under vehicle-only, infrastructure-only, and cooperative driving conditions within a unified multiple-choice question answering (MCQA) framework. The benchmark is organized into a twelve-task taxonomy spanning perception, prediction, and reasoning and planning, and is constructed through expert-verified MCQA annotation to enable fine-grained diagnosis of viewpoint-dependent capabilities. Benchmark results across ten representative state-of-the-art proprietary and open-source models show that viewpoint accessibility substantially affects performance, and infrastructure-side reasoning supports meaningful macroscopic traffic understanding. Results also indicate that cooperative reasoning remains challenging since it requires cross-view alignment and evidence integration rather than simply additional visual input. To address these challenges, we introduce V2X-MoE, a benchmark-aligned baseline with explicit view routing and viewpoint-specific LoRA experts. The strong performance of V2X-MoE further suggests that explicit viewpoint specialization is a promising direction for multi-view reasoning in autonomous driving. Overall, V2X-QA provides a foundation for studying multi-perspective reasoning, reliability, and cooperative physical intelligence in connected autonomous driving. The dataset and V2X-MoE resources are publicly available at: https://github.com/junwei0001/V2X-QA.

自动驾驶多模态协同推理评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。