用大模型生成罕见驾驶场景,提升自动驾驶安全性验证。
Generating Out-Of-Distribution Scenarios Using Language Models
- 用大模型构建分支树生成多样化异常驾驶场景。
- 在CARLA模拟器中按文本描述生成场景,多样性与异常度双评估。
- 验证语言模型对罕见场景的理解与导航能力,适合安全测试研究者。
自动驾驶系统部署需在多样真实环境中充分测试,应对边缘案例和分布外(OOD)场景,并通过全面安全验证确保系统在不可预测条件下的可靠运行。生成分布外驾驶场景对提升安全性至关重要,但因其长尾分布且在城市驾驶数据集中罕见,生成难度大。近期大型语言模型(LLMs)展现出零样本泛化和常识推理能力,为解决此问题带来新思路。本文提出一种框架,利用LLMs构建分支树,每条分支代表一个独特的OOD场景。这些场景通过自动化框架在CARLA模拟器中实现,场景增强与对应文本描述精准对齐。我们通过多样性指标评估场景丰富性,并引入新的“OOD程度”指标量化场景偏离常规城市驾驶的程度。此外,探索现代视觉-语言模型(VLMs)对模拟的OOD场景的解析与安全导航能力。结果为语言模型在城市驾驶分布外场景中的可靠性提供了重要洞见。
原文摘要 · Abstract (English)
The deployment of autonomous vehicles controlled by machine learning techniques requires extensive testing in diverse real-world environments, robust handling of edge cases and out-of-distribution scenarios, and comprehensive safety validation to ensure that these systems can navigate safely and effectively under unpredictable conditions. Addressing Out-Of-Distribution (OOD) driving scenarios is essential for enhancing safety, as OOD scenarios help validate the reliability of the models within the vehicle's autonomy stack. However, generating OOD scenarios is challenging due to their long-tailed distribution and rarity in urban driving dataset. Recently, Large Language Models (LLMs) have shown promise in autonomous driving, particularly for their zero-shot generalization and common-sense reasoning capabilities. In this paper, we leverage these LLM strengths to introduce a framework for generating diverse OOD driving scenarios. Our approach uses LLMs to construct a branching tree, where each branch represents a unique OOD scenario. These scenarios are then simulated in the CARLA simulator using an automated framework that aligns scene augmentation with the corresponding textual descriptions. We evaluate our framework through extensive simulations, and assess its performance via a diversity metric that measures the richness of the scenarios. Additionally, we introduce a new "OOD-ness" metric, which quantifies how much the generated scenarios deviate from typical urban driving conditions. Furthermore, we explore the capacity of modern Vision-Language Models (VLMs) to interpret and safely navigate through the simulated OOD scenarios. Our findings offer valuable insights into the reliability of language models in addressing OOD scenarios within the context of urban driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。