用大模型让自动驾驶更懂人,主动互动提升安全与效率
Large Language Model based Interactive Decision-Making for Autonomous Driving

- 用语义建模抽象交通场景,让大模型理解车辆间潜在意图
- 结合安全与效率约束,生成优化轨迹并通过自然语言沟通
- 决策过程透明可解释,适合复杂混合交通场景应用
在人类驾驶与自动驾驶混行的高冲突场景中,现有系统多采取过度保守策略,缺乏主动交互,导致公众接受度低。本文提出基于大语言模型的交互式决策框架,通过对象-过程方法论对多车场景进行语义建模,将底层感知数据抽象为对象、过程和关系,揭示隐含因果结构,辅助推理。在此基础上,大模型解析周围车辆的显性和隐性意图,在安全与效率双重约束下生成候选行驶行为。进一步通过蒙特卡洛采样生成扰动轨迹,评估并优化出可执行路径。最终,决策结果由大模型转化为简洁自然语言消息,经外部人机界面广播,实现从场景理解到行动再到语言的闭环。仿真测试表明,该方法在安全性、舒适性和效率上均优于传统基线;类图灵测试评估显示其决策具有高度类人特征。结果表明,语义场景抽象与大模型驱动的意图推理及语言化人机通信相结合,为密集混行交通中的交互式、可信自动驾驶提供了可行路径。
原文摘要 · Abstract (English)
In high-conflict mixed-traffic scenarios involving human-driven and autonomous vehicles, most existing autonomous driving systems default to overly conservative behaviors, lack proactive interaction, and consequently suffer from limited public acceptance. To mitigate intent misunderstandings and decision failures, we present a Large Language Model based interactive decision-making framework that augments scene understanding and intent-aware interaction to jointly improve safety and efficiency. The approach uses Object-Process Methodology to semantically model complex multi-vehicle scenes, abstracting low-level perceptual data into objects, processes, and relations, thereby streamlining reasoning over latent causal structure. Building on this representation, the Large Language Model parses both explicit and implicit intents of surrounding agents and, under jointly enforced safety and efficiency constraints, selects candidate maneuvers. We further generate perturbed trajectory candidates via Monte Carlo sampling and evaluate them to obtain an optimized executable trajectory. To foster transparency and coordination with nearby road users, the final decision is translated by the Large Language Model into concise natural-language messages and broadcast through an external Human-Machine Interface, completing a closed loop from scene understanding to action to language. Experiments in a cluster driving simulator demonstrate that the proposed method outperforms traditional baselines across safety, comfort, and efficiency metrics, while a Turing-test-style evaluation indicates a high degree of human-likeness in decision making. Besides, these results suggest that coupling semantic scene abstraction with Large Language Model mediated intent reasoning and language-based eHMI communication offers a practical pathway toward interactive, trustworthy autonomous driving in dense mixed traffic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。