让视觉语言模型读懂自动驾驶多视角场景,提升3D目标定位能力。
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
- 用可学习查询融合视觉语言模型与空间处理器,实现2D-3D对齐。
- 在150万张多视角图像上训练,3D视觉定位任务提升9.86%。
- 适合需要3D场景理解的自动驾驶研究者使用。
大型视觉语言模型(LVLM)显著提升了图像理解能力,其推理与理解能力为自动驾驶应用带来潜力。然而,现有研究多聚焦于前视视角和局部物体,难以实现全面场景理解。同时,现有LVLM缺乏2D与3D之间的映射关系,且3D目标定位与指令理解融合不足。为此,我们首次提出NuInteract,一个包含超过150万组多视角图像-语言对的大规模数据集,覆盖密集场景描述与多样化交互任务。进一步提出DriveMonkey框架,通过一系列可学习查询,将LVLM与空间处理器无缝集成。该空间处理器作为即插即用组件,可初始化预训练3D检测器以增强3D感知。实验表明,DriveMonkey优于通用LVLM,尤其在3D视觉定位任务中取得9.86%的显著提升。相关数据集与代码将在https://github.com/zc-zhao/DriveMonkey发布。
原文摘要 · Abstract (English)
The Large Visual-Language Models (LVLMs) have significantly advanced image understanding. Their comprehension and reasoning capabilities enable promising applications in autonomous driving scenarios. However, existing research typically focuses on front-view perspectives and partial objects within scenes, struggling to achieve comprehensive scene understanding. Meanwhile, existing LVLMs suffer from the lack of mapping relationship between 2D and 3D and insufficient integration of 3D object localization and instruction understanding. To tackle these limitations, we first introduce NuInteract, a large-scale dataset with over 1.5M multi-view image language pairs spanning dense scene captions and diverse interactive tasks. Furthermore, we propose DriveMonkey, a simple yet effective framework that seamlessly integrates LVLMs with a spatial processor using a series of learnable queries. The spatial processor, designed as a plug-and-play component, can be initialized with pre-trained 3D detectors to improve 3D perception. Our experiments show that DriveMonkey outperforms general LVLMs, especially achieving a 9.86% notable improvement on the 3D visual grounding task. The dataset and code will be released at https://github.com/zc-zhao/DriveMonkey.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。