让纯视觉自动驾驶模型同时懂语言、会开车、言行一致。
SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment
- 用视觉语言模型实现端到端的闭环驾驶与语言理解
- 在CARLA基准上达到顶尖驾驶表现,且语言任务结果优秀
- 无需激光雷达,适合追求低成本智能驾驶的团队
将大语言模型(LLMs)引入自动驾驶领域,旨在提升泛化能力与可解释性。然而,现有方法多侧重于驾驶或视觉-语言理解,难以兼顾两者。此外,主流的视觉问答方式在自动驾驶中仅在与动作空间对齐时才有意义,否则模型回答可能与其行为不一致。为此,我们提出SimLingo模型,能够同时处理三项任务:(1)闭环驾驶,(2)视觉-语言理解,(3)语言-动作对齐。该模型基于视觉语言模型(VLM),仅依赖摄像头,不使用激光雷达等昂贵传感器。SimLingo在广泛使用的CARLA模拟器上的Bench2Drive基准测试中取得领先性能,并成为CARLA挑战赛2024的获胜方案。同时,在多种语言相关任务上也展现出强劲表现,且保持了高水平的驾驶性能。
原文摘要 · Abstract (English)
Integrating large language models (LLMs) into autonomous driving has attracted significant attention with the hope of improving generalization and explainability. However, existing methods often focus on either driving or vision-language understanding but achieving both high driving performance and extensive language understanding remains challenging. In addition, the dominant approach to tackle vision-language understanding is using visual question answering. However, for autonomous driving, this is only useful if it is aligned with the action space. Otherwise, the model's answers could be inconsistent with its behavior. Therefore, we propose a model that can handle three different tasks: (1) closed-loop driving, (2) vision-language understanding, and (3) language-action alignment. Our model SimLingo is based on a vision language model (VLM) and works using only camera, excluding expensive sensors like LiDAR. SimLingo obtains state-of-the-art performance on the widely used CARLA simulator on the Bench2Drive benchmark and is the winning entry at the CARLA challenge 2024. Additionally, we achieve strong results in a wide variety of language-related tasks while maintaining high driving performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。