arXiv:2412.09043cs.CV2024-12NeurIPS被引 25

实时重建街景4D高斯模型,提升自动驾驶仿真精度

DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous Driving

  • 直接从环视视频预测4D高斯点,避免传统迭代耗时
  • 引入去重与动态静止解耦机制,显著提升重建质量
  • 适用于预训练、车辆适配与场景编辑,实用性强

真实感的街景4D重建对自动驾驶真实世界模拟器开发至关重要。然而,现有方法多为离线处理,依赖耗时的迭代过程,限制了实际应用。为此,我们提出大型4D高斯重建模型DrivingRecon,可直接从环视视频生成4D高斯点。为更好融合多视角图像,提出剪枝与扩张模块(PD-Block),消除相邻视图间重叠高斯点及冗余背景点。为增强跨时间信息建模,设计动态与静态特征解耦机制,更优学习几何与运动特征。实验表明,DrivingRecon在场景重建质量和新视角合成方面均优于现有方法。此外,我们探索其在模型预训练、车辆适配和场景编辑中的应用。代码已开源:https://github.com/EnVision-Research/DriveRecon。

原文摘要 · Abstract (English)

Photorealistic 4D reconstruction of street scenes is essential for developing real-world simulators in autonomous driving. However, most existing methods perform this task offline and rely on time-consuming iterative processes, limiting their practical applications. To this end, we introduce the Large 4D Gaussian Reconstruction Model (DrivingRecon), a generalizable driving scene reconstruction model, which directly predicts 4D Gaussian from surround view videos. To better integrate the surround-view images, the Prune and Dilate Block (PD-Block) is proposed to eliminate overlapping Gaussian points between adjacent views and remove redundant background points. To enhance cross-temporal information, dynamic and static decoupling is tailored to better learn geometry and motion features. Experimental results demonstrate that DrivingRecon significantly improves scene reconstruction quality and novel view synthesis compared to existing methods. Furthermore, we explore applications of DrivingRecon in model pre-training, vehicle adaptation, and scene editing. Our code is available at https://github.com/EnVision-Research/DriveRecon.

4D重建自动驾驶高斯渲染视觉建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。