用多模态鸟瞰图生成驾驶场景的自然语言描述,提升自动驾驶可解释性。
BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving
- 融合激光雷达与多视角图像,用新位置编码生成视点专属描述
- 10亿参数小模型在nuCaption上比现有方法高5%的BLEU得分
- 新增两个数据集,覆盖环境条件与物体定位,填补评测空白
自动驾驶技术有望重塑交通,但其广泛应用依赖于可解释、透明的决策系统。场景字幕任务通过生成对驾驶环境的自然语言描述,显著提升系统透明度、安全性和人机交互体验。本文提出BEV-LLM,一种轻量级3D场景字幕模型。该模型基于BEVFusion融合3D激光雷达点云与多视角图像,引入新颖的绝对位置编码以生成视点特定的场景描述。尽管仅采用10亿参数的基础模型,BEV-LLM在nuCaption数据集上表现优异,相比现有最先进方法最高提升5%的BLEU分数。此外,我们发布了两个新数据集:nuView(聚焦环境条件与视角)和GroundView(聚焦物体定位),用于更全面评估不同驾驶场景下的字幕生成能力,并提供了初步基准结果,验证了其有效性。
原文摘要 · Abstract (English)
Autonomous driving technology has the potential to transform transportation, but its wide adoption depends on the development of interpretable and transparent decision-making systems. Scene captioning, which generates natural language descriptions of the driving environment, plays a crucial role in enhancing transparency, safety, and human-AI interaction. We introduce BEV-LLM, a lightweight model for 3D captioning of autonomous driving scenes. BEV-LLM leverages BEVFusion to combine 3D LiDAR point clouds and multi-view images, incorporating a novel absolute positional encoding for view-specific scene descriptions. Despite using a small 1B parameter base model, BEV-LLM achieves competitive performance on the nuCaption dataset, surpassing state-of-the-art by up to 5\% in BLEU scores. Additionally, we release two new datasets - nuView (focused on environmental conditions and viewpoints) and GroundView (focused on object grounding) - to better assess scene captioning across diverse driving scenarios and address gaps in current benchmarks, along with initial benchmarking results demonstrating their effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。