arXiv:2506.00227cs.CVcs.AI2025-06被引 5

可控生成真实车祸视频,支持输入微调引发不同事故结果。

Ctrl-Crash: Controllable Diffusion for Realistic Car Crashes

  • 基于边界框、碰撞类型等信号控制车祸视频生成。
  • 在FVD、JEDi等指标上优于现有扩散模型。
  • 适合自动驾驶安全测试与事故复现场景。

近年来,视频扩散技术取得显著进展,但因驾驶数据集中事故事件稀缺,难以生成真实的车祸图像。提升交通安全性需要可现实且可控的事故模拟。为此,我们提出Ctrl-Crash,一种基于边界框、碰撞类型及初始图像帧等条件信号的可控车祸视频生成模型。该方法支持反事实场景生成,输入微小变化即可导致截然不同的碰撞结果。为实现推理时细粒度控制,我们采用独立可调缩放系数的无分类器引导机制。在定量(如FVD、JEDi)和定性(人类评估物理真实性和视频质量)指标上,相比先前基于扩散的方法,均达到最先进性能。

原文摘要 · Abstract (English)

Video diffusion techniques have advanced significantly in recent years; however, they struggle to generate realistic imagery of car crashes due to the scarcity of accident events in most driving datasets. Improving traffic safety requires realistic and controllable accident simulations. To tackle the problem, we propose Ctrl-Crash, a controllable car crash video generation model that conditions on signals such as bounding boxes, crash types, and an initial image frame. Our approach enables counterfactual scenario generation where minor variations in input can lead to dramatically different crash outcomes. To support fine-grained control at inference time, we leverage classifier-free guidance with independently tunable scales for each conditioning signal. Ctrl-Crash achieves state-of-the-art performance across quantitative video quality metrics (e.g., FVD and JEDi) and qualitative measurements based on a human-evaluation of physical realism and video quality compared to prior diffusion-based methods.

视频生成扩散模型车祸模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。