arXiv:2607.19701cs.CV2026-07

用扩散模型生成高危交通场景,提升视觉语言模型自动驾驶的安全评估能力。

SafeGen: Goal-Conditioned Video Diffusion of Safety-Critical Scenarios for VLM-Based Autonomous Driving

论文配图:SafeGen: Goal-Conditioned Video Diffusion of Safety-Critical Scenarios for VLM-Based Autonomous Driving
图 1 · 摘自论文原文
  • 以灾难性结局为目标,通过条件扩散生成连贯视频轨迹。
  • 利用视觉语言模型推断潜在风险,生成结构化高危场景指令。
  • 可显著提升自动驾驶系统安全评估分数,适合安全测试与模型优化。

视觉语言模型(VLM)在自动驾驶系统中日益普及,亟需对罕见但高危的交通场景进行严格安全评估。其中,与弱势道路使用者的交互是真实世界事故的主要来源。现有方法多依赖模拟器生成场景,存在显著的仿真到现实差距,难以捕捉真实、多样且意外的人车互动动态。本文提出SafeGen,一种面向视觉语言模型自动驾驶(VLMAD)的高危场景生成框架。核心思想是将场景生成建模为以预设灾难性终态为目标的条件扩散过程,利用该强监督信号引导生成时间连贯的视频轨迹,自然演化至高危状态。在此基础上,提出上下文感知终态推理(Context Grounded End State Reasoning),借助VLM分析正常驾驶上下文,推断人车交互中的潜在脆弱性,生成结构化的终态规范以诱发高风险场景。进一步提出终态条件视频演化(End State Conditioned Video Evolution),将语义威胁转化为物理合理的视觉动态:通过深度感知几何投影在场景中实例化高危主体,并采用边界约束扩散生成运动一致、时序连贯的中间帧。在3个VLMAD系统上的实验表明,与最先进基线相比,SafeGen平均使基于VLM判官的总体评分(Judge Overall Score)提升24.25%。此外,对VLMAD进行微调后,在真实驾驶场景中的性能平均提升15.9%。

原文摘要 · Abstract (English)

VLMs are increasingly deployed in AD systems, creating an urgent need for rigorous safety evaluation under rare yet safety-critical scenarios. Among these, interactions with vulnerable road users represent a major source of real-world failures. However, existing safety-critical scenario generation methods predominantly rely on simulator-based pipelines, which suffer from a substantial sim-to-real gap and often fail to capture realistic, diverse, and unforeseen human-vehicle interaction dynamics. We present SafeGen, a goal-conditioned diffusion framework for safety-critical scenario generation in VLMADs. Our key insight is to formulate scenario generation as a goal-conditioned diffusion process, where a predefined catastrophic end-state serves as a strong supervisory signal, guiding the generation of temporally coherent video trajectories that naturally evolve toward safety-critical outcomes. Building on this formulation, we introduce Context Grounded End State Reasoning, which leverages VLMs to analyze benign driving contexts and infer latent vulnerabilities in human-vehicle interactions, producing structured end-state specifications that induce high-risk scenarios. Conditioned on these targets, we further propose End State Conditioned Video Evolution, which grounds semantic threats into physically plausible visual dynamics. Specifically, we instantiate high-risk agents within the scene via depth-aware geometric projection, followed by boundary-conditioned diffusion to generate intermediate frames with consistent motion patterns and temporal coherence. Extensive experiments across 3 VLMADs demonstrate that SafeGen increases the Judge Overall Score, a metric using a VLM judge to evaluate VLMADs' understanding and decision-making, by 24.25% on average compared to SoTA baselines. Furthermore, fine-tuning a VLMAD improves performance in real-world driving scenes by an average of 15.9%.

自动驾驶扩散模型安全评估视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。