用大模型模拟人类变道行为,发现其能复现部分驾驶习惯但安全表现差异大。
General-purpose LLMs as Models of Human Driver Behavior: The Case of Simplified Merging

- 将大模型嵌入简化变道场景,作为闭环驾驶员代理运行
- 模型对空间线索响应类似人类,但动态速度响应不一致,安全表现差异显著
- 提示词设计影响模型行为,不同模型需定制化优化
人类行为模型在自动驾驶虚拟安全评估中至关重要,但现有模型常在可解释性与灵活性间权衡。通用大语言模型(LLMs)提供新可能:单一模型无需参数调优即可部署于多种场景。然而,其对人类驾驶行为的捕捉能力尚不明确。本文将两个通用大模型(OpenAI o3 和 Google Gemini 2.5 Pro)作为独立闭环驾驶员代理,嵌入简化一维变道场景,通过定量与定性分析对比其行为与人类数据。两者均再现了人类间歇性操作控制及对空间线索的战术依赖,但未能一致捕捉对动态速度线索的响应,且安全性能差异显著。系统性提示词消融实验表明,提示组件构成模型特异性归纳偏置,无法跨模型迁移。结果表明,通用大模型或可作为即插即用的人类行为模型用于自动驾驶评估流程,但未来研究需更深入理解其失效模式并验证其有效性。
原文摘要 · Abstract (English)
Human behavior models are essential as behavior references and for simulating human agents in virtual safety assessment of automated vehicles (AVs), yet current models face a trade-off between interpretability and flexibility. General-purpose large language models (LLMs) offer a promising alternative: a single model potentially deployable without parameter fitting across diverse scenarios. However, what LLMs can and cannot capture about human driving behavior remains poorly understood. We address this gap by embedding two general-purpose LLMs (OpenAI o3 and Google Gemini 2.5 Pro) as standalone, closed-loop driver agents in a simplified one-dimensional merging scenario and comparing their behavior against human data using quantitative and qualitative analyses. Both models reproduce human-like intermittent operational control and tactical dependencies on spatial cues. However, neither consistently captures the human response to dynamic velocity cues, and safety performance diverges sharply between models. A systematic prompt ablation study reveals that prompt components act as model-specific inductive biases that do not transfer across LLMs. These findings suggest that general-purpose LLMs could potentially serve as standalone, ready-to-use human behavior models in AV evaluation pipelines, but future research is needed to better understand their failure modes and ensure their validity as models of human driving behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。