arXiv:2409.09575cs.RO2024-09中稿 · CVPR被引 22

用大模型生成自然语言描述的交通场景,提升自动驾驶测试真实性

Traffic Scene Generation from Natural Language Description for Autonomous Vehicles with Large Language Model

  • 分步式框架将文本转为可行驶道路与多智能体行为
  • 碰撞率低至3.5%,生成场景更安全可控
  • 适合自动驾驶系统测试与行为推理训练

从自然语言生成真实且可控的交通场景,可显著提升自动驾驶系统的研发与评估效率。然而该任务面临三大挑战:(1) 将自由文本映射为空间有效且语义一致的布局;(2) 在无预设位置的情况下构建场景;(3) 规划多智能体行为并选择符合配置的车道。为此,我们提出模块化框架TTSG,包括提示分析、道路检索、智能体规划及新颖的计划感知道路排序算法。虽使用大语言模型(LLMs)作为通用规划器,但设计上将其嵌入严格控制的流水线,确保结构、可行性与场景多样性。特别地,排序策略保障了智能体行为与道路几何的一致性,实现无需预设路径或起始点的场景生成。框架支持常规与高危场景,以及多阶段事件组合。在SafeBench上的实验表明,本方法在三个关键场景中平均碰撞率最低(3.5%)。此外,基于生成场景训练的驾驶字幕模型,在动作推理上提升超过30 CIDEr分。结果验证了该框架在灵活性、可解释性与安全性方面的优势。

原文摘要 · Abstract (English)

Generating realistic and controllable traffic scenes from natural language can greatly enhance the development and evaluation of autonomous driving systems. However, this task poses unique challenges: (1) grounding free-form text into spatially valid and semantically coherent layouts, (2) composing scenarios without predefined locations, and (3) planning multi-agent behaviors and selecting roads that respect agents' configurations. To address these, we propose a modular framework, TTSG, comprising prompt analysis, road retrieval, agent planning, and a novel plan-aware road ranking algorithm to solve these challenges. While large language models (LLMs) are used as general planners, our design integrates them into a tightly controlled pipeline that enforces structure, feasibility, and scene diversity. Notably, our ranking strategy ensures consistency between agent actions and road geometry, enabling scene generation without predefined routes or spawn points. The framework supports both routine and safety-critical scenarios, as well as multi-stage event composition. Experiments on SafeBench demonstrate that our method achieves the lowest average collision rate (3.5\%) across three critical scenarios. Moreover, driving captioning models trained on our generated scenes improve action reasoning by over 30 CIDEr points. These results underscore our proposed framework for flexible, interpretable, and safety-oriented simulation.

交通生成大模型自动驾驶场景模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。