用因果实体生成驾驶事故视频,提升自动驾驶安全测试能力
Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis
- 基于事故原因与驾驶员注视点定位关键参与者和行为
- 在154万帧数据上训练,生成视频的因果敏感度超越现有模型
- 适合自动驾驶安全验证、事故复现与视频生成研究者使用
以第一人称视角理解汽车事故的原因与影响,对自动驾驶安全至关重要。合成反映因果关系的事故视频有助于测试系统应对罕见事故的能力。然而,如何将真实视频中的因果关系融入生成视频仍具挑战。本文认为准确识别事故参与者及其相关行为至关重要。为此,提出新型扩散模型Causal-VidSyn,利用事故原因描述和驾驶员注视点信息,通过事故原因问答与凝视条件选择模块实现因果实体定位。为支持该模型,构建了目前最大的驾驶事故场景下驾驶员注视数据集Drive-Gaze(含154万帧注视数据)。大量实验表明,Causal-VidSyn在帧质量与因果敏感度方面均优于现有视频扩散模型,在事故视频编辑、正常到事故视频生成及文本到视频生成任务中表现优异。
原文摘要 · Abstract (English)
Egocentricly comprehending the causes and effects of car accidents is crucial for the safety of self-driving cars, and synthesizing causal-entity reflected accident videos can facilitate the capability test to respond to unaffordable accidents in reality. However, incorporating causal relations as seen in real-world videos into synthetic videos remains challenging. This work argues that precisely identifying the accident participants and capturing their related behaviors are of critical importance. In this regard, we propose a novel diffusion model, Causal-VidSyn, for synthesizing egocentric traffic accident videos. To enable causal entity grounding in video diffusion, Causal-VidSyn leverages the cause descriptions and driver fixations to identify the accident participants and behaviors, facilitated by accident reason answering and gaze-conditioned selection modules. To support Causal-VidSyn, we further construct Drive-Gaze, the largest driver gaze dataset (with 1.54M frames of fixations) in driving accident scenarios. Extensive experiments show that Causal-VidSyn surpasses state-of-the-art video diffusion models in terms of frame quality and causal sensitivity in various tasks, including accident video editing, normal-to-accident video diffusion, and text-to-video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。