用视觉语言模型提升自动驾驶感知与规划,实现复杂场景端到端行驶
AppleVLM: End-to-end Autonomous Driving with Advanced Perception and Planning-Enhanced Vision-Language Models
- 引入时空融合视觉编码器和显式规划编码器,增强多视角感知与决策鲁棒性
- 在CARLA双基准测试中达顶尖表现,真实AGV平台验证复杂户外环境可行性
- 通过分层思维链微调,缓解语言指令偏差,适合高阶自动驾驶系统研发者
端到端自动驾驶将感知、决策与控制整合于统一学习框架,具有广阔前景。近期,视觉语言模型(VLMs)因其在多样化未见场景中提升模型鲁棒性与泛化能力的潜力而受到关注。然而,现有基于VLM的方法仍存在车道感知不佳、语言理解偏差及难处理极端情况等问题。为此,我们提出AppleVLM,一种增强感知与规划的VLM模型,用于实现稳健的端到端驾驶。AppleVLM引入新型视觉编码器与规划策略编码器:首先,视觉编码器利用可变形变压器机制融合多时序多视角图像的时空信息,提升对摄像头差异的鲁棒性,并支持跨车辆平台扩展部署;其次,不同于传统方法,AppleVLM引入专用规划模态,编码显式的鸟瞰图空间信息,降低导航指令中的语言偏差;最后,通过分层思维链微调的VLM解码器,融合视觉、语言与规划特征,输出鲁棒的驾驶航点。我们在两个CARLA基准上进行闭环实验,达到当前最优驾驶性能;此外,将AppleVLM部署于AGV平台,在复杂室外环境中成功实现真实世界端到端自动驾驶。
原文摘要 · Abstract (English)
End-to-end autonomous driving has emerged as a promising paradigm integrating perception, decision-making, and control within a unified learning framework. Recently, Vision-Language Models (VLMs) have gained significant attention for their potential to enhance the robustness and generalization of end-to-end driving models in diverse and unseen scenarios. However, existing VLM-based approaches still face challenges, including suboptimal lane perception, language understanding biases, and difficulties in handling corner cases. To address these issues, we propose AppleVLM, an advanced perception and planning-enhanced VLM model for robust end-to-end driving. AppleVLM introduces a novel vision encoder and a planning strategy encoder to improve perception and decision-making. Firstly, the vision encoder fuses spatial-temporal information from multi-view images across multiple timesteps using a deformable transformer mechanism, enhancing robustness to camera variations and facilitating scalable deployment across different vehicle platforms. Secondly, unlike traditional VLM-based approaches, AppleVLM introduces a dedicated planning modality that encodes explicit Bird's-Eye-View spatial information, mitigating language biases in navigation instructions. Finally, a VLM decoder fine-tuned by a hierarchical Chain-of-Thought integrates vision, language, and planning features to output robust driving waypoints. We evaluate AppleVLM in closed-loop experiments on two CARLA benchmarks, achieving state-of-the-art driving performance. Furthermore, we deploy AppleVLM on an AGV platform and successfully showcase real-world end-to-end autonomous driving in complex outdoor environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。