首个面向4D LiDAR的多模态大模型基准,支持时空理解任务。
B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding
- 构建可扩展的4D LiDAR数据生成流水线,实现真实场景建模。
- 提出首个直接处理原始4D LiDAR的多模态大模型,实现时空推理。
- 适用于自动驾驶、机器人感知等动态环境理解场景。
理解动态室外环境需捕捉复杂物体交互及其随时间演变。基于激光雷达的4D点云提供精确空间几何与丰富时间线索,是表征真实场景的理想方式。然而,由于缺乏高质量、模态特异的标注以及适配其高维结构的多模态大模型架构,4D激光雷达在多模态大模型(MLLM)中的应用仍不充分。为此,我们提出B4DL,一个专为训练和评估MLLM在4D激光雷达理解上设计的新基准。同时,我们开发了可扩展的数据生成管道和一个首次直接处理原始4D激光雷达的MLLM模型,通过将其与语言理解相连接实现端到端融合。结合我们的数据集与基准,该模型为动态室外环境中的时空推理提供了统一解决方案。我们公开了渲染的4D激光雷达视频、生成数据集及多样化场景下的推理输出:https://github.com/ccho4702/B4DL。
原文摘要 · Abstract (English)
Understanding dynamic outdoor environments requires capturing complex object interactions and their evolution over time. LiDAR-based 4D point clouds provide precise spatial geometry and rich temporal cues, making them ideal for representing real-world scenes. However, despite their potential, 4D LiDAR remains underexplored in the context of Multimodal Large Language Models (MLLMs) due to the absence of high-quality, modality-specific annotations and the lack of MLLM architectures capable of processing its high-dimensional composition. To address these challenges, we introduce B4DL, a new benchmark specifically designed for training and evaluating MLLMs on 4D LiDAR understanding. In addition, we propose a scalable data generation pipeline and an MLLM model that, for the first time, directly processes raw 4D LiDAR by bridging it with language understanding. Combined with our dataset and benchmark, our model offers a unified solution for spatio-temporal reasoning in dynamic outdoor environments. We provide rendered 4D LiDAR videos, generated dataset, and inference outputs on diverse scenarios at: https://github.com/ccho4702/B4DL
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。