用大模型提升自动驾驶的感知与决策能力,探索落地路径。
LLM4AD: Large Language Models for Autonomous Driving -- Concept, Review, Benchmark, Experiments, and Future Trends
- 构建专用大模型框架LLM4AD,融合语言理解与驾驶场景推理
- 推出三类评测基准,在仿真与真实车辆上验证性能
- 适合研究自动驾驶与大模型交叉的学者及工程团队
随着大语言模型(LLMs)的广泛应用与快速发展,将其应用于自动驾驶技术的兴趣和需求日益增长。凭借其自然语言理解和推理能力,LLMs有望提升自动驾驶系统在感知、场景理解及交互式决策等环节的表现。本文首次提出面向自动驾驶的大语言模型(LLM4AD)概念,并综述现有相关研究。随后,构建了全面的评估基准,包括用于指令遵循与推理能力测试的LaMPilot-Bench、CARLA Leaderboard 1.0仿真基准以及多视角视觉问答任务的NuPlanQA。此外,在真实自动驾驶平台上开展了大量实验,考察云端与边缘端部署下个性化决策与运动控制的效果。进一步探讨了将语言扩散模型融入自动驾驶的未来趋势,提出ViLaD(Vision-Language Diffusion)框架。最后,分析了LLM4AD面临的主要挑战,涵盖延迟、部署、安全隐私、安全性、可信任性与透明度,以及个性化等问题。
原文摘要 · Abstract (English)
With the broader adoption and highly successful development of Large Language Models (LLMs), there has been growing interest and demand for applying LLMs to autonomous driving technology. Driven by their natural language understanding and reasoning capabilities, LLMs have the potential to enhance various aspects of autonomous driving systems, from perception and scene understanding to interactive decision-making. This paper first introduces the novel concept of designing Large Language Models for Autonomous Driving (LLM4AD), followed by a review of existing LLM4AD studies. Then, a comprehensive benchmark is proposed for evaluating the instruction-following and reasoning abilities of LLM4AD systems, which includes LaMPilot-Bench, CARLA Leaderboard 1.0 Benchmark in simulation and NuPlanQA for multi-view visual question answering. Furthermore, extensive real-world experiments are conducted on autonomous vehicle platforms, examining both on-cloud and on-edge LLM deployment for personalized decision-making and motion control. Next, the future trends of integrating language diffusion models into autonomous driving are explored, exemplified by the proposed ViLaD (Vision-Language Diffusion) framework. Finally, the main challenges of LLM4AD are discussed, including latency, deployment, security and privacy, safety, trust and transparency, and personalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。