用自然语言生成真实互动的自动驾驶测试场景,让指令与结果更贴合。
LinguaSim: Interactive Multi-Vehicle Testing Scenario Generation via Natural Language Instruction Based on Large Language Models
- 基于大模型将自然语言转为可交互的3D驾驶场景。
- 生成场景的危险度与描述一致,危急情况响应时间仅0.072秒。
- 适合自动驾驶安全测试与训练,尤其关注意图精准匹配的场景设计。
自动驾驶测试与训练场景生成受到广泛关注。尽管大语言模型(LLMs)推动了新方法发展,但现有方法难以兼顾指令遵循准确率与真实道路环境的拟真度。为降低描述复杂性,当前方法常将场景限制在二维或开环仿真中,背景车辆行为预设且不互动。我们提出LinguaSim,一个基于大模型的框架,能将自然语言转化为具有动态车辆交互的高保真3D场景,确保输入描述与生成场景高度一致。通过反馈校准模块进一步提升生成精度,增强对用户意图的还原度。该框架通过结合场景描述和自动驾驶模型引导,约束对抗性车辆行为。实验表明,LinguaSim能生成与自然语言描述相匹配的多级危急场景(危险描述响应时间0.072秒,安全描述3.532秒;舒适度评分0.654 vs. 0.764),其优化模块有效抑制初始输出中的过度激进行为,使碰撞率从46.9%降至6.3%,更贴近用户意图。
原文摘要 · Abstract (English)
The generation of testing and training scenarios for autonomous vehicles has drawn significant attention. While Large Language Models (LLMs) have enabled new scenario generation methods, current methods struggle to balance command adherence accuracy with the realism of real-world driving environments. To reduce scenario description complexity, these methods often compromise realism by limiting scenarios to 2D, or open-loop simulations where background vehicles follow predefined, non-interactive behaviors. We propose LinguaSim, an LLM-based framework that converts natural language into realistic, interactive 3D scenarios, ensuring both dynamic vehicle interactions and faithful alignment between the input descriptions and the generated scenarios. A feedback calibration module further refines the generation precision, improving fidelity to user intent. By bridging the gap between natural language and closed-loop, interactive simulations, LinguaSim constrains adversarial vehicle behaviors using both the scenario description and the autonomous driving model guiding them. This framework facilitates the creation of high-fidelity scenarios that enhance safety testing and training. Experiments show LinguaSim can generate scenarios with varying criticality aligned with different natural language descriptions (ACT: 0.072 s for dangerous vs. 3.532 s for safe descriptions; comfortability: 0.654 vs. 0.764), and its refinement module effectively reduces excessive aggressiveness in LinguaSim's initial outputs, lowering the crash rate from 46.9% to 6.3% to better match user intentions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。