arXiv:2603.11380cs.CV2026-03中稿 · CVPR

构建多模态驾驶问答数据集,提升自动驾驶对恶劣场景的理解能力

DriveXQA: Cross-modal Visual Question Answering for Adverse Driving Scene Understanding

  • 设计双交叉注意力融合模块,高效整合多传感器视觉信息
  • 在雾天等恶劣条件下,模型性能显著优于基线(GPTScore 53.5 vs 25.1)
  • 适合自动驾驶感知与多模态理解方向的研究者参考

融合互补模态的传感器对于稳定、全面理解异常驾驶场景至关重要。然而,多模态大语言模型(MLLMs)在利用多传感器信息理解自动驾驶中的恶劣驾驶场景方面仍处于探索阶段。为此,我们提出DriveXQA,一个面向自动驾驶视觉问答的多模态数据集。该数据集包含四种视觉模态、五种传感器故障情况和五种天气条件,共102,505个问答对,按全局场景级、非中心级和自车中心级分为三类。由于现有MLLM框架未采用多种互补视觉模态作为输入,我们设计了MVX-LLM,一种具有双交叉注意力(DCA)投影器的轻量级架构,以缓解信息冗余。实验表明,在雾天等挑战性条件下,所提DCA方法性能显著提升(GPTScore:53.5 vs. 25.1)。

原文摘要 · Abstract (English)

Fusing sensors with complementary modalities is crucial for maintaining a stable and comprehensive understanding of abnormal driving scenes. However, Multimodal Large Language Models (MLLMs) are underexplored for leveraging multi-sensor information to understand adverse driving scenarios in autonomous vehicles. To address this gap, we propose the DriveXQA, a multimodal dataset for autonomous driving VQA. In addition to four visual modalities, five sensor failure cases, and five weather conditions, it includes $102,505$ QA pairs categorized into three types: global scene level, allocentric level, and ego-vehicle centric level. Since no existing MLLM framework adopts multiple complementary visual modalities as input, we design MVX-LLM, a token-efficient architecture with a Dual Cross-Attention (DCA) projector that fuses the modalities to alleviate information redundancy. Experiments demonstrate that our DCA achieves improved performance under challenging conditions such as foggy (GPTScore: $53.5$ vs. $25.1$ for the baseline).

自动驾驶多模态视觉问答感知融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。