arXiv:2410.13571cs.CV2024-10CVPR被引 139

用世界模型生成4D驾驶场景,让仿真更真实、动态更连贯。

DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene Representation

  • 用世界模型当数据机器,生成带时空一致性的新轨迹视频。
  • 相比现有方法,生成质量提升32.1%~46.4%(FID),动态一致性提高15.6%~43.5%。
  • 适合自动驾驶仿真、4D重建研究者,尤其关注复杂变道等场景的团队。

闭环仿真对推进端到端自动驾驶系统至关重要。当前传感器仿真方法如NeRF和3DGS主要依赖与训练数据分布高度一致的条件,多局限于前向驾驶场景,难以渲染复杂操作(如变道、加速、减速)。近期自动驾驶世界模型虽展示了生成多样化驾驶视频的潜力,但仍限于2D视频生成,缺乏动态驾驶环境所需的时空一致性。本文提出DriveDreamer4D,利用世界模型先验增强4D驾驶场景表征。具体而言,将世界模型作为数据机器合成新轨迹视频,显式利用结构化条件控制交通元素的时空一致性。同时提出亲缘数据训练策略,促进真实与合成数据融合以优化4DGS。据我们所知,DriveDreamer4D是首个使用视频生成模型改进驾驶场景4D重建的工作。实验表明,其在新轨迹视角下的生成质量显著提升,相较PVG、S3Gaussian和Deformable-GS,FID分别降低32.1%、46.4%和16.3%;驾驶主体的时空一致性也大幅提升,用户评估及NTA-IoU指标相对提升22.6%、43.5%和15.6%。

原文摘要 · Abstract (English)

Closed-loop simulation is essential for advancing end-to-end autonomous driving systems. Contemporary sensor simulation methods, such as NeRF and 3DGS, rely predominantly on conditions closely aligned with training data distributions, which are largely confined to forward-driving scenarios. Consequently, these methods face limitations when rendering complex maneuvers (e.g., lane change, acceleration, deceleration). Recent advancements in autonomous-driving world models have demonstrated the potential to generate diverse driving videos. However, these approaches remain constrained to 2D video generation, inherently lacking the spatiotemporal coherence required to capture intricacies of dynamic driving environments. In this paper, we introduce DriveDreamer4D, which enhances 4D driving scene representation leveraging world model priors. Specifically, we utilize the world model as a data machine to synthesize novel trajectory videos, where structured conditions are explicitly leveraged to control the spatial-temporal consistency of traffic elements. Besides, the cousin data training strategy is proposed to facilitate merging real and synthetic data for optimizing 4DGS. To our knowledge, DriveDreamer4D is the first to utilize video generation models for improving 4D reconstruction in driving scenarios. Experimental results reveal that DriveDreamer4D significantly enhances generation quality under novel trajectory views, achieving a relative improvement in FID by 32.1%, 46.4%, and 16.3% compared to PVG, S3Gaussian, and Deformable-GS. Moreover, DriveDreamer4D markedly enhances the spatiotemporal coherence of driving agents, which is verified by a comprehensive user study and the relative increases of 22.6%, 43.5%, and 15.6% in the NTA-IoU metric.

4D重建世界模型自动驾驶视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。