arXiv:2503.03074cs.ROcs.CV2025-03被引 32

用地图特征让大模型更懂自动驾驶指令,提升闭环驾驶表现。

BEVDriver: Leveraging BEV Maps in LLMs for Robust Closed-Loop Driving

  • 将鸟瞰图特征融合进大模型,实现视觉与语言联合感知
  • 在LangAuto基准上驾驶得分比顶尖方法高18.9%
  • 适合关注大模型决策与自动驾驶结合的研究者

自动驾驶有望推动未来高效出行,但需通过安全、可靠和透明的驾驶建立信任。大型语言模型(LLMs)具备推理与自然语言理解能力,可作为通用决策者进行自主车辆路径规划,并与人类交互、适应人类设计的环境。然而,现有方法难以融合三维空间定位与大模型的推理能力。本文提出BEVDriver,一种基于大模型的端到端闭环驾驶系统,使用潜在鸟瞰图(BEV)特征作为感知输入。该模型包含一个BEV编码器,可高效处理多视角图像与3D激光雷达点云;在统一潜在空间中,通过Q-Former对齐语言指令,将特征传入大模型以预测并规划精确轨迹,同时考虑导航指令与关键场景。在LangAuto基准测试中,本模型驾驶得分较最先进方法最高提升18.9%。

原文摘要 · Abstract (English)

Autonomous driving has the potential to set the stage for more efficient future mobility, requiring the research domain to establish trust through safe, reliable and transparent driving. Large Language Models (LLMs) possess reasoning capabilities and natural language understanding, presenting the potential to serve as generalized decision-makers for ego-motion planning that can interact with humans and navigate environments designed for human drivers. While this research avenue is promising, current autonomous driving approaches are challenged by combining 3D spatial grounding and the reasoning and language capabilities of LLMs. We introduce BEVDriver, an LLM-based model for end-to-end closed-loop driving in CARLA that utilizes latent BEV features as perception input. BEVDriver includes a BEV encoder to efficiently process multi-view images and 3D LiDAR point clouds. Within a common latent space, the BEV features are propagated through a Q-Former to align with natural language instructions and passed to the LLM that predicts and plans precise future trajectories while considering navigation instructions and critical scenarios. On the LangAuto benchmark, our model reaches up to 18.9% higher performance on the Driving Score compared to SoTA methods.

自动驾驶大模型鸟瞰图闭环控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。