用激光雷达增强视觉语言模型,让自动驾驶更懂3D空间
Spatial-aware Vision Language Model for Autonomous Driving
- 引入激光雷达点云,逐步融合到预训练视觉语言模型中
- 在多个驾驶基准上显著提升空间感知与决策可靠性
- 专为自动驾驶设计,适合追求高安全性的系统开发者
尽管视觉语言模型(VLMs)通过利用语言模型中的常识,在端到端自动驾驶中展现出巨大潜力,但其依赖2D图像线索进行复杂场景理解与决策,严重制约了安全性与可靠性。现有基于图像的方法难以实现精确的度量空间推理与几何推断,导致驾驶策略不可靠。为此,我们提出LVLDrive(LiDAR-Vision-Language),一种专门用于将现有VLMs升级为具备稳健3D度量空间理解能力的框架,通过引入激光雷达点云作为额外模态。核心挑战在于缓解异构3D数据对预训练VLM造成的灾难性干扰。为此,我们设计渐进式融合Q-Former,逐步注入激光雷达特征,确保VLM原有知识库的稳定与保留。此外,我们构建了面向空间感知的问题回答(SA-QA)数据集,显式训练模型掌握高级3D感知与推理能力。在多个驾驶基准上的大量实验表明,相较于纯视觉方法,LVLDrive在场景理解、度量空间感知和可靠驾驶决策方面均表现更优。本工作强调了明确引入3D度量数据对于构建可信的VLM驱动自动驾驶系统的重要性。
原文摘要 · Abstract (English)
While Vision-Language Models (VLMs) show significant promise for end-to-end autonomous driving by leveraging the common sense embedded in language models, their reliance on 2D image cues for complex scene understanding and decision-making presents a critical bottleneck for safety and reliability. Current image-based methods struggle with accurate metric spatial reasoning and geometric inference, leading to unreliable driving policies. To bridge this gap, we propose LVLDrive (LiDAR-Vision-Language), a novel framework specifically designed to upgrade existing VLMs with robust 3D metric spatial understanding for autonomous driving by incoperating LiDAR point cloud as an extra input modality. A key challenge lies in mitigating the catastrophic disturbance introduced by disparate 3D data to the pre-trained VLMs. To this end, we introduce a Gradual Fusion Q-Former that incrementally injects LiDAR features, ensuring the stability and preservation of the VLM's existing knowledge base. Furthermore, we develop a spatial-aware question-answering (SA-QA) dataset to explicitly teach the model advanced 3D perception and reasoning capabilities. Extensive experiments on driving benchmarks demonstrate that LVLDrive achieves superior performance compared to vision-only counterparts across scene understanding, metric spatial perception, and reliable driving decision-making. Our work highlights the necessity of explicit 3D metric data for building trustworthy VLM-based autonomous systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。