arXiv:2512.02844cs.ROcs.LG2025-12

用视觉语言模型指导扩散模型,自动生成逼真且危险的自动驾驶测试场景。

VLM as Strategist: Adaptive Generation of Safety-critical Testing Scenarios via Guided Diffusion

  • 利用VLM理解场景风险,生成测试目标和引导指令。
  • 实现对仿真中背景车辆的实时精准控制,支持闭环交互。
  • 适合自动驾驶安全验证团队,提升长尾极端场景生成效率。

自动驾驶系统(ADS)的安全部署依赖于全面的测试评估。然而,能有效暴露系统漏洞的安全关键场景在现实世界中极为稀少。现有场景生成方法难以高效构建具有高保真度、高危性和强互动性的长尾场景,尤其缺乏对被测车辆(VUT)的实时动态响应能力。本文提出一种融合视觉语言模型(VLM)高层语义理解与自适应引导扩散模型细粒度生成能力的安全关键测试场景生成框架。该框架采用三层分层架构:战略层由VLM决定场景生成目标,战术层制定引导函数,操作层执行引导扩散。首先构建高质量基础扩散模型以学习真实驾驶数据分布;随后设计自适应引导扩散方法,实现闭环仿真中对背景车辆(BVs)的实时精确控制。VLM通过深度场景理解与风险推理,自主生成场景目标与引导函数,最终驱动扩散模型完成由VLM主导的场景生成。实验表明,所提方法可高效生成真实、多样且高度交互的安全关键场景。案例研究进一步验证了方法的适应性与由VLM引导的生成性能。

原文摘要 · Abstract (English)

The safe deployment of autonomous driving systems (ADSs) relies on comprehensive testing and evaluation. However, safety-critical scenarios that can effectively expose system vulnerabilities are extremely sparse in the real world. Existing scenario generation methods face challenges in efficiently constructing long-tail scenarios that ensure fidelity, criticality, and interactivity, while particularly lacking real-time dynamic response capabilities to the vehicle under test (VUT). To address these challenges, this paper proposes a safety-critical testing scenario generation framework that integrates the high-level semantic understanding capabilities of Vision Language Models (VLMs) with the fine-grained generation capabilities of adaptive guided diffusion models. The framework establishes a three-layer hierarchical architecture comprising a strategic layer for VLM-directed scenario generation objective determination, a tactical layer for guidance function formulation, and an operational layer for guided diffusion execution. We first establish a high-quality fundamental diffusion model that learns the data distribution of real driving scenarios. Next, we design an adaptive guided diffusion method that enables real-time, precise control of background vehicles (BVs) in closed-loop simulation. The VLM is then incorporated to autonomously generate scenario generation objectives and guidance functions through deep scenario understanding and risk reasoning, ultimately guiding the diffusion model to achieve VLM-directed scenario generation. Experimental results demonstrate that the proposed method can efficiently generate realistic, diverse, and highly interactive safety-critical testing scenarios. Furthermore, case studies validate the adaptability and VLM-directed generation performance of the proposed method.

自动驾驶场景生成VLM扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。