用大模型增强自动驾驶系统,让机器更懂复杂路况和行车意图。
LeAD: The LLM Enhanced Planning System Converged with End-to-end Autonomous Driving
- 双速率架构:高频端到端模型保实时,低频大模型做语义推理。
- 在CARLA仿真中完成率93%,排行榜得分71分,优于传统方案。
- 适合研究自动驾驶语义理解与决策的学者及工程师参考。
城市自动驾驶大规模部署的主要障碍在于复杂场景和边缘情况频发。现有系统难以有效解析交通语境中的语义信息,也难以判断其他交通参与者意图,导致决策偏离熟练驾驶员的思维模式。我们提出LeAD,一种融合基于模仿学习的端到端(E2E)框架与大语言模型(LLM)增强的双速率自动驾驶架构。高频端到端子系统维持实时感知-规划-控制循环,低频LLM模块通过多模态感知融合高精地图,利用思维链(CoT)推理在基础规划器能力受限时生成最优决策。在CARLA模拟器上的实验表明,LeAD能更好应对非典型场景,在Leaderboard V1基准上获得71分,路线完成率达93%。
原文摘要 · Abstract (English)
A principal barrier to large-scale deployment of urban autonomous driving systems lies in the prevalence of complex scenarios and edge cases. Existing systems fail to effectively interpret semantic information within traffic contexts and discern intentions of other participants, consequently generating decisions misaligned with skilled drivers' reasoning patterns. We present LeAD, a dual-rate autonomous driving architecture integrating imitation learning-based end-to-end (E2E) frameworks with large language model (LLM) augmentation. The high-frequency E2E subsystem maintains real-time perception-planning-control cycles, while the low-frequency LLM module enhances scenario comprehension through multi-modal perception fusion with HD maps and derives optimal decisions via chain-of-thought (CoT) reasoning when baseline planners encounter capability limitations. Our experimental evaluation in the CARLA Simulator demonstrates LeAD's superior handling of unconventional scenarios, achieving 71 points on Leaderboard V1 benchmark, with a route completion of 93%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。