用自然语言生成对抗性驾驶场景,提升自动驾驶测试可控性。
LD-Scene: LLM-Guided Diffusion for Controllable Generation of Adversarial Safety-Critical Driving Scenarios
- 结合大模型与扩散模型,通过语言指令生成对抗性场景
- 在nuScenes数据集上生成真实、多样且有效的对抗场景
- 支持细粒度控制,适合自动驾驶安全测试人员使用
确保自动驾驶系统的安全性和鲁棒性需要在关键安全场景下进行全面评估。然而,这些关键安全场景在真实驾驶数据中极为稀少且难以获取,给自动驾驶性能评估带来重大挑战。现有方法通常缺乏可控性和用户友好性,需依赖大量专家知识。为此,我们提出LD-Scene,一种将大语言模型(LLMs)与潜在扩散模型(LDMs)结合的新框架,通过自然语言实现用户可控的对抗性场景生成。该方法包含一个捕捉真实驾驶轨迹分布的LDM,以及一个将用户查询转化为对抗损失函数的LLM引导模块,支持生成与用户需求一致的场景。引导模块集成基于思维链(CoT)的代码生成器和代码调试器,增强引导函数生成的可控性与鲁棒性。在nuScenes数据集上的大量实验表明,LD-Scene在生成真实、多样且高效的对抗场景方面达到当前最优性能,并可对对抗行为进行细粒度控制,从而实现针对特定驾驶场景的有效测试。
原文摘要 · Abstract (English)
Ensuring the safety and robustness of autonomous driving systems necessitates a comprehensive evaluation in safety-critical scenarios. However, these safety-critical scenarios are rare and difficult to collect from real-world driving data, posing significant challenges to effectively assessing the performance of autonomous vehicles. Typical existing methods often suffer from limited controllability and lack user-friendliness, as extensive expert knowledge is essentially required. To address these challenges, we propose LD-Scene, a novel framework that integrates Large Language Models (LLMs) with Latent Diffusion Models (LDMs) for user-controllable adversarial scenario generation through natural language. Our approach comprises an LDM that captures realistic driving trajectory distributions and an LLM-based guidance module that translates user queries into adversarial loss functions, facilitating the generation of scenarios aligned with user queries. The guidance module integrates an LLM-based Chain-of-Thought (CoT) code generator and an LLM-based code debugger, enhancing the controllability and robustness in generating guidance functions. Extensive experiments conducted on the nuScenes dataset demonstrate that LD-Scene achieves state-of-the-art performance in generating realistic, diverse, and effective adversarial scenarios. Furthermore, our framework provides fine-grained control over adversarial behaviors, thereby facilitating more effective testing tailored to specific driving scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。