用世界模型生成逼真交通场景,提升自动驾驶事故预判能力
World Model-Based End-to-End Scene Generation for Accident Anticipation in Autonomous Driving
- 用领域提示引导的世界模型生成高分辨率驾驶场景
- 在新数据集上事故预判准确率显著提升,提前时间更长
- 适合自动驾驶安全系统研发者和算法工程师参考
可靠的交通事故预判对推动自动驾驶系统至关重要。然而,这一目标受限于两大挑战:高质量、多样化的训练数据稀缺,以及环境干扰或传感器缺陷导致的关键物体线索缺失。为此,我们提出一个融合生成场景增强与自适应时序推理的综合框架。具体而言,我们构建了一个视频生成流水线,利用受领域知识引导的世界模型生成高分辨率、统计一致的驾驶场景,尤其丰富了边缘案例和复杂交互的覆盖。同时,我们设计了一种动态预测模型,通过强化图卷积与空洞时序算子编码时空关系,有效应对数据不完整和瞬时视觉噪声问题。此外,我们发布了一个新基准数据集,以更好捕捉真实世界的多样化驾驶风险。在公开及新发布的数据集上进行的大量实验表明,该框架提升了事故预判的准确率与提前时间,为当前自动驾驶安全应用中的数据与建模局限提供了稳健解决方案。
原文摘要 · Abstract (English)
Reliable anticipation of traffic accidents is essential for advancing autonomous driving systems. However, this objective is limited by two fundamental challenges: the scarcity of diverse, high-quality training data and the frequent absence of crucial object-level cues due to environmental disruptions or sensor deficiencies. To tackle these issues, we propose a comprehensive framework combining generative scene augmentation with adaptive temporal reasoning. Specifically, we develop a video generation pipeline that utilizes a world model guided by domain-informed prompts to create high-resolution, statistically consistent driving scenarios, particularly enriching the coverage of edge cases and complex interactions. In parallel, we construct a dynamic prediction model that encodes spatio-temporal relationships through strengthened graph convolutions and dilated temporal operators, effectively addressing data incompleteness and transient visual noise. Furthermore, we release a new benchmark dataset designed to better capture diverse real-world driving risks. Extensive experiments on public and newly released datasets confirm that our framework enhances both the accuracy and lead time of accident anticipation, offering a robust solution to current data and modeling limitations in safety-critical autonomous driving applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。