让自动驾驶车听懂乘客的自由指令并安全执行
Open-Ended Instruction Realization with LLM-Enabled Multi-Planner Scheduling in Autonomous Vehicles
- 用大模型理解乘客指令,调度多个规划器生成动作
- 任务完成率显著高于基线,且对大模型延迟容忍度高
- 适合研究人机交互与自动驾驶融合的学者
当前人机交互研究普遍忽视自动驾驶中乘客的变道需求。自然语言虽是直观接口,但将乘客的开放性指令转化为控制信号,同时保持可解释性和可追溯性仍具挑战。本文提出一种基于大语言模型(LLM)的指令实现框架:由LLM解析指令,生成调度多个基于模型预测控制(MPC)的运动规划器的可执行脚本,并根据实时反馈动态调整,最终将规划轨迹转换为控制信号。该调度式设计在不同时间尺度上解耦语义推理与车辆控制,构建从高层指令到底层动作的透明决策链。由于缺乏高保真评估工具,本文引入闭环设置下的开放指令实现基准。大量实验表明,该框架显著提升任务完成率,降低LLM查询成本,安全性与合规性达到专用自动驾驶方法水平,并具备较强抗大模型推理延迟能力。
原文摘要 · Abstract (English)
Most Human-Machine Interaction (HMI) research overlooks the maneuvering needs of passengers in autonomous driving (AD). Natural language offers an intuitive interface, yet translating passenger open-ended instructions into control signals, without sacrificing interpretability and traceability, remains a challenge. This study proposes an instruction-realization framework that leverages a large language model (LLM) to interpret instructions, generates executable scripts that schedule multiple model predictive control (MPC)-based motion planners based on real-time feedback, and converts planned trajectories into control signals. This scheduling-centric design decouples semantic reasoning from vehicle control at different timescales, establishing a transparent, traceable decision-making chain from high-level instructions to low-level actions. Due to the absence of high-fidelity evaluation tools, this study introduces a benchmark for open-ended instruction realization in a closed-loop setting. Comprehensive experiments reveal that the framework significantly improves task-completion rates over instruction-realization baselines, reduces LLM query costs, achieves safety and compliance on par with specialized AD approaches, and exhibits considerable tolerance to LLM inference latency. For more qualitative illustrations and a clearer understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。