arXiv:2409.06450cs.ROcs.AI2024-09被引 31

用大模型生成真实多样的自动驾驶测试场景,提升测试效率与泛化能力。

Multimodal Large Language Model Driven Scenario Testing for Autonomous Vehicles

  • 基于多模态大模型生成可控制的测试场景,融合世界知识与推理能力。
  • 在三种复杂场景中验证生成结果的逼真性与多样性,支持事故报告还原。
  • 结合检索增强与自优化机制,提升场景生成的准确性和实用性,适合自动驾驶研发者。

corner case的生成已成为自动驾驶汽车在上路前高效测试的关键环节。然而,现有方法难以满足多样化的测试需求,且缺乏对未见情境的泛化能力,降低了生成场景的可用性。为此,我们提出OmniTester:一种基于多模态大语言模型(LLM)的框架,充分挖掘LLM的广泛世界知识与推理能力,用于在仿真环境中生成真实且多样的测试场景,为自动驾驶系统测试与评估提供稳健解决方案。除提示工程外,我们采用Simulation of Urban Mobility(SUMO)工具简化LLM生成代码的复杂度;同时引入检索增强生成(RAG)与自改进机制,提升LLM对场景的理解力,从而增强生成现实场景的能力。实验表明,本方法在生成三类复杂挑战性场景方面具备良好可控性与真实性,并展示了基于LLM泛化能力重构事故报告中描述新场景的有效性。

原文摘要 · Abstract (English)

The generation of corner cases has become increasingly crucial for efficiently testing autonomous vehicles prior to road deployment. However, existing methods struggle to accommodate diverse testing requirements and often lack the ability to generalize to unseen situations, thereby reducing the convenience and usability of the generated scenarios. A method that facilitates easily controllable scenario generation for efficient autonomous vehicles (AV) testing with realistic and challenging situations is greatly needed. To address this, we proposed OmniTester: a multimodal Large Language Model (LLM) based framework that fully leverages the extensive world knowledge and reasoning capabilities of LLMs. OmniTester is designed to generate realistic and diverse scenarios within a simulation environment, offering a robust solution for testing and evaluating AVs. In addition to prompt engineering, we employ tools from Simulation of Urban Mobility to simplify the complexity of codes generated by LLMs. Furthermore, we incorporate Retrieval-Augmented Generation and a self-improvement mechanism to enhance the LLM's understanding of scenarios, thereby increasing its ability to produce more realistic scenes. In the experiments, we demonstrated the controllability and realism of our approaches in generating three types of challenging and complex scenarios. Additionally, we showcased its effectiveness in reconstructing new scenarios described in crash report, driven by the generalization capability of LLMs.

自动驾驶大模型场景生成仿真测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。