让自动驾驶模型理解360度全景空间关系,回答更复杂的交通场景问题。
BeLLA: End-to-End Birds Eye View Large Language Assistant for Autonomous Driving
- 用统一的360度鸟瞰图表示融合多摄像头信息,增强空间感知
- 在需要空间推理的问题上提升9.3%准确率,尤其擅长判断物体相对位置
- 适合做智能驾驶决策解释、复杂场景问答的开发者和研究者
视觉-语言模型(VLM)和多模态语言模型(MLLM)的快速发展推动了自动驾驶领域对场景理解、上下文推理和可解释决策的进步。然而,现有方法多依赖单视角编码器,无法充分利用多摄像头系统的空间结构,或使用聚合的多视角特征,缺乏统一的空间表示,难以进行以自身为中心的方向判断、物体关系分析和整体环境理解。为此,我们提出BeLLA,一种端到端架构,将统一的360°鸟瞰图(BEV)表示与大型语言模型结合,实现自动驾驶场景下的问答任务。我们在NuScenes-QA和DriveLM两个基准上评估,BeLLA在需强空间推理的问题(如相对位置、邻近物体行为理解)上持续领先,某些任务绝对提升达9.3%;在其他类别中也表现竞争力,展现出对多样化问题的处理能力。
原文摘要 · Abstract (English)
The rapid development of Vision-Language models (VLMs) and Multimodal Language Models (MLLMs) in autonomous driving research has significantly reshaped the landscape by enabling richer scene understanding, context-aware reasoning, and more interpretable decision-making. However, a lot of existing work often relies on either single-view encoders that fail to exploit the spatial structure of multi-camera systems or operate on aggregated multi-view features, which lack a unified spatial representation, making it more challenging to reason about ego-centric directions, object relations, and the wider context. We thus present BeLLA, an end-to-end architecture that connects unified 360° BEV representations with a large language model for question answering in autonomous driving. We primarily evaluate our work using two benchmarks - NuScenes-QA and DriveLM, where BeLLA consistently outperforms existing approaches on questions that require greater spatial reasoning, such as those involving relative object positioning and behavioral understanding of nearby objects, achieving up to +9.3% absolute improvement in certain tasks. In other categories, BeLLA performs competitively, demonstrating the capability of handling a diverse range of questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。