arXiv:2607.19528cs.CVcs.AI2026-07中稿 · IEEE IV 2026被引 1

用大模型理解三维动态交通场景,提升自动驾驶安全判断能力。

D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models

论文配图:D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models
图 1 · 摘自论文原文
  • 统一处理2D视频与3D时序数据,构建轻量一体化架构
  • 在KITTI QA上比基线提升11%准确率,验证3D数据价值
  • 新推出WaymoQA数据集,评估复杂路况下的多模态理解

多模态大语言模型(MLLM)的进展推动了自动驾驶端到端系统的开发,但现有研究主要聚焦于2D图像与视频。本文关注3D传感器(如激光雷达与双目相机)在MLLM中的应用,解决激光雷达数据稀疏、非网格结构带来的融合难题。提出D3VL框架,首次实现2D与3D时序数据的统一建模。该模型旨在回答交通场景理解与安全相关问题,在KITTI问答数据集上较基线提升11%性能。同时引入扩展的Waymo QA数据集,评估模型在多样化驾驶条件下的3D时序理解能力。代码与数据集可在补充网站获取:https://automotivesafety-lvlm.github.io

原文摘要 · Abstract (English)

Recent advances in Multimodal Large Language Models (MLLMs) have triggered the development of end-to-end MLLMs for autonomous driving. However, the main emphasis to date has been for MLLMs using 2D images and videos. In contrast, this paper considers MLLM effectiveness using 3D sensors, particularly LiDAR and stereo cameras. LiDAR presents unique challenges to integration within an MLLM, largely because of data sparsity and lack of a grid structure for the data. For similar reasons, fusion of camera and LiDAR data within an MLLM pipeline is also uncommon. However, most autonomous systems rely on LiDAR-based sensing, and incorporating 3D data has been proven to improve performance in traditional 3D scene perception tasks. This paper presents D3VL, a novel MLLM framework that integrates 2D and 3D time-series data in a single but simple architecture. The model aims to answer questions involving traffic scene understanding and safety. D3VL shows an 11% improvement in the KITTI Question-Answering (QA) dataset compared to baseline methods in processing 2D and 3D time-series data. This paper further introduces the Waymo QA dataset extension, which assesses models' capabilities in processing 3D and time-series data under diverse driving conditions. D3VL implementation code and WaymoQA extension can be found on our supplemental website: https://automotivesafety-lvlm.github.io

自动驾驶多模态大模型3D感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。