arXiv:2505.17727cs.CV2025-05被引 6

生成真实世界多视角危险驾驶视频,用于测试自动驾驶系统可靠性。

SafeMVDrive: Multi-view Safety-Critical Driving Video Synthesis in the Real World Domain

  • 结合视觉上下文与大模型,生成更真实的危险驾驶轨迹。
  • 通过分阶段可控轨迹生成,确保碰撞规避场景真实可信。
  • 采用扩散模型合成高质量多视角视频,适合自动驾驶系统压力测试。

安全关键场景在自动驾驶中虽罕见却至关重要,现有方法仅能生成轨迹、仿真或单视角视频,难以满足端到端自动驾驶系统对真实世界多视角视频数据的需求。为此,我们提出SafeMVDrive,首个基于真实世界域的多视角危险驾驶视频生成框架。该框架将安全关键轨迹生成器与先进多视角视频生成器有机结合:首先,通过引入视觉上下文并利用GRPO微调的视觉-语言模型,增强轨迹生成器的场景理解能力,生成更真实、上下文感知的轨迹;其次,针对现有视频生成器难以真实呈现碰撞事件的问题,设计两阶段可控轨迹生成机制,生成具备避碰能力的轨迹,保障视频质量和安全性;最后,采用基于扩散模型的多视角视频生成器,从生成轨迹合成高质量的安全关键驾驶视频。在端到端自动驾驶规划器上的实验表明,使用生成数据后碰撞率显著提升,验证了该框架在压力测试规划模块方面的有效性。代码、示例及数据集已公开:https://zhoujiawei3.github.io/SafeMVDrive/。

原文摘要 · Abstract (English)

Safety-critical scenarios are rare yet pivotal for evaluating and enhancing the robustness of autonomous driving systems. While existing methods generate safety-critical driving trajectories, simulations, or single-view videos, they fall short of meeting the demands of advanced end-to-end autonomous systems (E2E AD), which require real-world, multi-view video data. To bridge this gap, we introduce SafeMVDrive, the first framework designed to generate high-quality, safety-critical, multi-view driving videos grounded in real-world domains. SafeMVDrive strategically integrates a safety-critical trajectory generator with an advanced multi-view video generator. To tackle the challenges inherent in this integration, we first enhance scene understanding ability of the trajectory generator by incorporating visual context -- which is previously unavailable to such generator -- and leveraging a GRPO-finetuned vision-language model to achieve more realistic and context-aware trajectory generation. Second, recognizing that existing multi-view video generators struggle to render realistic collision events, we introduce a two-stage, controllable trajectory generation mechanism that produces collision-evasion trajectories, ensuring both video quality and safety-critical fidelity. Finally, we employ a diffusion-based multi-view video generator to synthesize high-quality safety-critical driving videos from the generated trajectories. Experiments conducted on an E2E AD planner demonstrate a significant increase in collision rate when tested with our generated data, validating the effectiveness of SafeMVDrive in stress-testing planning modules. Our code, examples, and datasets are publicly available at: https://zhoujiawei3.github.io/SafeMVDrive/.

自动驾驶视频生成扩散模型安全测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。