arXiv:2505.00156cs.CV2025-05CVPR被引 9

将视觉语言模型与3D信息结合,提升自动驾驶场景理解能力

V3LMA: Visual 3D-enhanced Language Model for Autonomous Driving

  • 用检测结果和视频生成文本,增强3D场景理解
  • 在LingoQA上取得0.56分,无需微调即提升性能
  • 适合关注自动驾驶感知与决策的开发者

大型视觉语言模型(LVLM)在多个领域展现出强大的视觉场景理解能力。但在自动驾驶场景中,其对3D环境的感知有限,影响了对动态交通环境的全面安全理解。为此,我们提出V3LMA,一种将大语言模型(LLM)与LVLM融合的新方法,通过从目标检测和视频输入生成文本描述,显著提升性能且无需微调。借助专门的预处理流程提取3D物体数据,该方法增强了复杂交通场景中的情境意识与决策能力,在LingoQA基准测试中达到0.56分。我们进一步探索了多种融合策略与标记组合,旨在推动交通场景的更优解析,最终实现更安全的自动驾驶系统。

原文摘要 · Abstract (English)

Large Vision Language Models (LVLMs) have shown strong capabilities in understanding and analyzing visual scenes across various domains. However, in the context of autonomous driving, their limited comprehension of 3D environments restricts their effectiveness in achieving a complete and safe understanding of dynamic surroundings. To address this, we introduce V3LMA, a novel approach that enhances 3D scene understanding by integrating Large Language Models (LLMs) with LVLMs. V3LMA leverages textual descriptions generated from object detections and video inputs, significantly boosting performance without requiring fine-tuning. Through a dedicated preprocessing pipeline that extracts 3D object data, our method improves situational awareness and decision-making in complex traffic scenarios, achieving a score of 0.56 on the LingoQA benchmark. We further explore different fusion strategies and token combinations with the goal of advancing the interpretation of traffic scenes, ultimately enabling safer autonomous driving systems.

自动驾驶视觉语言模型3D理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。