arXiv:2601.05806cs.RO2026-01被引 1

让自动驾驶系统听懂人话,用大模型实现自然语言交互

Modular Autonomy with Conversational Interaction: An LLM-driven Framework for Decision Making in Autonomous Driving

  • 用大模型将人话转为驾驶指令,支持五类交互操作
  • 系统在仿真中成功执行全部指令,响应时间高效
  • 适合自动驾驶人机交互、安全可控系统设计的研究者

大型语言模型(LLMs)为自动驾驶系统(ADS)提供了自然语言接口的新可能,突破传统固定输入的限制。本文针对人类语言复杂性与模块化自动驾驶软件动作空间之间的映射难题,提出一种融合LLM交互层与Autoware开源系统的框架。该系统使乘客可通过高阶指令进行操作,包括查询状态信息和修改驾驶行为。方法基于三大核心组件:交互类别分类体系、面向应用的领域特定语言(DSL)用于指令转换、以及保障安全的验证层。采用两阶段LLM架构,通过明确执行状态提供可解释反馈,提升透明度。评估表明系统具备良好时效性与翻译鲁棒性,仿真验证了全部五类交互指令的成功执行。本工作为可扩展、基于DSL的安全自动驾驶交互提供了基础。

原文摘要 · Abstract (English)

Recent advancements in Large Language Models (LLMs) offer new opportunities to create natural language interfaces for Autonomous Driving Systems (ADSs), moving beyond rigid inputs. This paper addresses the challenge of mapping the complexity of human language to the structured action space of modular ADS software. We propose a framework that integrates an LLM-based interaction layer with Autoware, a widely used open-source software. This system enables passengers to issue high-level commands, from querying status information to modifying driving behavior. Our methodology is grounded in three key components: a taxonomization of interaction categories, an application-centric Domain Specific Language (DSL) for command translation, and a safety-preserving validation layer. A two-stage LLM architecture ensures high transparency by providing feedback based on the definitive execution status. Evaluation confirms the system's timing efficiency and translation robustness. Simulation successfully validated command execution across all five interaction categories. This work provides a foundation for extensible, DSL-assisted interaction in modular and safety-conscious autonomy stacks.

自动驾驶大模型人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。