arXiv:2608.20129cs.MAcs.CL2026-08

用大模型推理增强自动驾驶决策,兼顾安全与场景理解。

Multi-Agent Orchestration with the Common-Sense Reasoning Capabilities of LLMs for Autonomous Driving

论文配图:Multi-Agent Orchestration with the Common-Sense Reasoning Capabilities of LLMs for Autonomous Driving
图 1 · 摘自论文原文
  • 用大模型做常识推理,协调强化学习与传统控制
  • 在随机化CARLA场景中实现高成功率的动态驾驶
  • 适合关注智能决策与系统融合的研究者

自动驾驶车辆需具备鲁棒的感知与决策能力以应对多样且未见的场景。尽管强化学习和基于规则的方法可提供有效控制与安全保障,但在需要上下文推理的情境下性能可能下降。大型语言模型(LLMs)在理解多模态信息和生成上下文推理方面表现出强大能力,但直接用于车辆控制会引入延迟和幻觉风险。为此,提出一种混合框架:通过调度器协调经PPO训练的强化学习与PID控制,并在整个系统中应用LLM的常识推理。此外,利用大模型迭代优化强化学习奖励函数,以适应动态驾驶环境。该框架在多样化环境与交通条件下,于高度随机化的CARLA场景中进行了评估。结果表明,将基于大模型的推理与传统自动驾驶方法结合,可在保持结构化控制与安全机制的同时,提升系统表现。

原文摘要 · Abstract (English)

Autonomous vehicles require robust perception and decision-making capabilities to operate in diverse and unseen scenarios. While reinforcement learning and rule-based methods can provide effective control and safety mechanisms, their performance may degrade in situations requiring contextual reasoning. Large Language Models (LLMs) have demonstrated strong capabilities in understanding multimodal information and generating contextual reasoning, however, their use for direct vehicle control can introduce latency and hallucination risks. To address these limitations, a hybrid framework is proposed. This system uses an orchestrator to coordinate PPO-trained reinforcement learning and PID control, with LLM common-sense reasoning applied throughout the framework. LLM reasoning is further employed iteratively to refine the RL reward function for dynamic driving environments. The proposed framework is evaluated in highly randomized CARLA scenarios under diverse environmental and traffic conditions. The results demonstrate the potential of integrating LLM-based reasoning with conventional autonomous driving methods while retaining structured control and safety mechanism.

自动驾驶大模型多智能体强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。