用大模型提升自动驾驶的决策能力,解决模块化与端到端的痛点
A Survey on Large Language Model-empowered Autonomous Driving
- 将大语言模型融入自动驾驶模块或端到端系统,增强推理与理解能力
- 大模型可缓解模块间目标不一致问题,提升复杂场景应对能力
- 适合关注智能驾驶前沿、大模型应用的研究者与工程师
人工智能在自动驾驶研究中起关键作用,推动其向智能化与高效化发展。当前自动驾驶技术主要沿两条路径:模块化与端到端。模块化将驾驶任务分解为感知、预测、规划和控制等模块,分别训练,但因各模块训练目标不一致,集成效果存在偏差;端到端尝试通过单一模型直接从传感器数据映射到控制信号,但受限于综合特征学习能力,难以应对长尾事件与复杂城市交通场景。面对双重挑战,研究者认为具备强大推理能力与广泛知识理解的大语言模型(LLMs)可能成为突破口,有望赋予自动驾驶系统更深层次的理解与决策能力。本文全面分析了大语言模型在自动驾驶系统中的潜在应用,探讨其在模块化与端到端架构中的优化策略,重点研究其如何解决现有方案中的问题。同时讨论一个核心问题:基于大语言模型的人工通用智能(AGI)能否成为实现高等级自动驾驶的关键?进一步剖析大语言模型在推动自动驾驶技术发展中可能面临的局限与挑战。
原文摘要 · Abstract (English)
Artificial intelligence (AI) plays a crucial role in autonomous driving (AD) research, propelling its development towards intelligence and efficiency. Currently, the development of AD technology follows two main technical paths: modularization and end-to-end. Modularization decompose the driving task into modules such as perception, prediction, planning, and control, and train them separately. Due to the inconsistency of training objectives between modules, the integrated effect suffers from bias. End-to-end attempts to address this issue by utilizing a single model that directly maps from sensor data to control signals. This path has limited learning capabilities in a comprehensive set of features and struggles to handle unpredictable long-tail events and complex urban traffic scenarios. In the face of challenges encountered in both paths, many researchers believe that large language models (LLMs) with powerful reasoning capabilities and extensive knowledge understanding may be the solution, expecting LLMs to provide AD systems with deeper levels of understanding and decision-making capabilities. In light of the challenges faced by both paths, many researchers believe that LLMs, with their powerful reasoning abilities and extensive knowledge, could offer a solution. To understand if LLMs could enhance AD, this paper conducts a thorough analysis of the potential applications of LLMs in AD systems, including exploring their optimization strategies in both modular and end-to-end approaches, with a particular focus on how LLMs can tackle the problems and challenges present in current solutions. Furthermore, we discuss an important question: Can LLM-based artificial general intelligence (AGI) be a key to achieve high-level AD? We further analyze the potential limitations and challenges that LLMs may encounter in promoting the development of AD technology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。