用生成模型合成带3D一致性的动态驾驶场景,适合自驾车训练。
DreamDrive: Generative 4D Scene Modeling from Street View Images
- 结合生成与重建优势,用扩散模型生成视觉参考并转为4D场景
- 在nuScenes和街景数据上实现高保真、3D一致的新视角视频生成
- 无需标注即可自动分离静态动态元素,适合真实道路场景
从车载行驶轨迹合成逼真的视觉观测是大规模自驾车模型训练的关键。重建类方法依赖昂贵的物体标注,难以泛化到真实道路;生成类模型虽更通用,但常缺乏3D视觉一致性。本文提出DreamDrive,一种融合生成与重建优势的4D时空场景生成方法,可合成具有3D一致性的动态驾驶视频。我们利用视频扩散模型生成一系列视觉参考,再通过新型混合高斯表示将其提升至4D,并基于高斯点渲染生成3D一致的驾驶视频。生成先验使方法能在真实道路数据上生成高质量4D场景,神经渲染保障3D一致性。在nuScenes和街景图像上的实验表明,DreamDrive能生成可控且泛化的4D驾驶场景,以高保真度合成新视角驾驶视频,自监督分解静态与动态元素,并提升自动驾驶感知与规划任务性能。
原文摘要 · Abstract (English)
Synthesizing photo-realistic visual observations from an ego vehicle's driving trajectory is a critical step towards scalable training of self-driving models. Reconstruction-based methods create 3D scenes from driving logs and synthesize geometry-consistent driving videos through neural rendering, but their dependence on costly object annotations limits their ability to generalize to in-the-wild driving scenarios. On the other hand, generative models can synthesize action-conditioned driving videos in a more generalizable way but often struggle with maintaining 3D visual consistency. In this paper, we present DreamDrive, a 4D spatial-temporal scene generation approach that combines the merits of generation and reconstruction, to synthesize generalizable 4D driving scenes and dynamic driving videos with 3D consistency. Specifically, we leverage the generative power of video diffusion models to synthesize a sequence of visual references and further elevate them to 4D with a novel hybrid Gaussian representation. Given a driving trajectory, we then render 3D-consistent driving videos via Gaussian splatting. The use of generative priors allows our method to produce high-quality 4D scenes from in-the-wild driving data, while neural rendering ensures 3D-consistent video generation from the 4D scenes. Extensive experiments on nuScenes and street view images demonstrate that DreamDrive can generate controllable and generalizable 4D driving scenes, synthesize novel views of driving videos with high fidelity and 3D consistency, decompose static and dynamic elements in a self-supervised manner, and enhance perception and planning tasks for autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。