arXiv:2504.00379cs.CV2025-04CVPR被引 21

用视觉标记替代文字坐标,提升自动驾驶问答的空间理解能力。

MPDrive: Improving Spatial Understanding with Marker-Based Prompt Learning for Autonomous Driving

  • 用标记图像替代文本坐标,实现视觉与语言的空间表达一致
  • 在DriveLM和CODA-LM上达到当前最优性能,尤其擅长复杂空间推理
  • 适合需要精准空间感知的自动驾驶视觉问答研究者使用

自动驾驶视觉问答(AD-VQA)旨在基于驾驶场景图像回答感知、预测和规划相关问题,高度依赖模型的空间理解能力。以往方法通常通过文本形式表达坐标,导致视觉坐标与文字描述之间存在语义鸿沟,阻碍空间信息的准确传递并增加表达负担。为此,我们提出一种新型基于标记的提示学习框架MPDrive,用简洁的视觉标记表示空间坐标,确保语言表达的一致性,提升AD-VQA中视觉感知与空间表达的准确性。具体地,我们利用检测专家在物体区域叠加数字标签生成标记图像,将复杂的文本坐标生成转化为直接的文本-视觉标记预测。此外,我们将原始图像与标记图像融合为场景级特征,并结合检测先验获得实例级特征。通过整合这些特征,构建双粒度视觉提示,激发大语言模型的空间感知能力。在DriveLM和CODA-LM数据集上的大量实验表明,MPDrive在需要复杂空间理解的任务中表现卓越,达到当前最优水平。

原文摘要 · Abstract (English)

Autonomous driving visual question answering (AD-VQA) aims to answer questions related to perception, prediction, and planning based on given driving scene images, heavily relying on the model's spatial understanding capabilities. Prior works typically express spatial information through textual representations of coordinates, resulting in semantic gaps between visual coordinate representations and textual descriptions. This oversight hinders the accurate transmission of spatial information and increases the expressive burden. To address this, we propose a novel Marker-based Prompt learning framework (MPDrive), which represents spatial coordinates by concise visual markers, ensuring linguistic expressive consistency and enhancing the accuracy of both visual perception and spatial expression in AD-VQA. Specifically, we create marker images by employing a detection expert to overlay object regions with numerical labels, converting complex textual coordinate generation into straightforward text-based visual marker predictions. Moreover, we fuse original and marker images as scene-level features and integrate them with detection priors to derive instance-level features. By combining these features, we construct dual-granularity visual prompts that stimulate the LLM's spatial perception capabilities. Extensive experiments on the DriveLM and CODA-LM datasets show that MPDrive achieves state-of-the-art performance, particularly in cases requiring sophisticated spatial understanding.

自动驾驶视觉问答空间理解提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。