让自动驾驶视频生成更符合物理规律,提升真实感与感知任务表现。
Physical Informed Driving World Model
- 通过坐标对齐、3D运动引导和框坐标指导,融合物理约束生成视频。
- 在Nuscenes上实现3.96 FID和38.06 FVD,优于现有方法。
- 适合需要高真实感训练数据的自动驾驶视觉研究者。
自动驾驶依赖高质量、大规模多视角驾驶视频训练感知模型,用于3D目标检测、分割和轨迹预测。尽管世界模型能低成本生成逼真驾驶视频,但如何确保视频遵循基本物理规律仍具挑战,包括相对与绝对运动、遮挡关系、空间一致性及时间一致性。为此,我们提出DrivePhysica,通过三项创新实现:(1) 坐标系统对齐模块,融合相对与绝对运动特征以增强运动理解;(2) 实例流引导模块,通过高效3D光流提取保障时序一致性;(3) 框坐标引导模块,改善空间关系建模并准确解析遮挡层级。基于物理原则建模,我们在Nuscenes数据集上达到3.96 FID和38.06 FVD的生成质量,显著提升下游感知任务性能。
原文摘要 · Abstract (English)
Autonomous driving requires robust perception models trained on high-quality, large-scale multi-view driving videos for tasks like 3D object detection, segmentation and trajectory prediction. While world models provide a cost-effective solution for generating realistic driving videos, challenges remain in ensuring these videos adhere to fundamental physical principles, such as relative and absolute motion, spatial relationship like occlusion and spatial consistency, and temporal consistency. To address these, we propose DrivePhysica, an innovative model designed to generate realistic multi-view driving videos that accurately adhere to essential physical principles through three key advancements: (1) a Coordinate System Aligner module that integrates relative and absolute motion features to enhance motion interpretation, (2) an Instance Flow Guidance module that ensures precise temporal consistency via efficient 3D flow extraction, and (3) a Box Coordinate Guidance module that improves spatial relationship understanding and accurately resolves occlusion hierarchies. Grounded in physical principles, we achieve state-of-the-art performance in driving video generation quality (3.96 FID and 38.06 FVD on the Nuscenes dataset) and downstream perception tasks. Our project homepage: https://metadrivescape.github.io/papers_project/DrivePhysica/page.html
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。