arXiv:2409.12753cs.CV2024-09AAAI被引 61

无需精确标定即可实时重建驾驶场景,支持多视角灵活输入。

DrivingForward: Feed-forward 3D Gaussian Splatting for Driving Scene Reconstruction from Flexible Surround-view Input

  • 自监督联合训练姿态、深度与高斯粒子网络,无需真实深度和相机参数。
  • 单帧独立预测高斯参数,实现多帧输入下的前向推理重建。
  • 在nuScenes上优于现有方法,适合自动驾驶场景建模应用。

我们提出DrivingForward,一种从灵活多视角输入中重建驾驶场景的前馈式高斯点渲染模型。车载摄像头获取的驾驶场景图像通常稀疏且重叠有限,车辆运动导致相机外参难以准确获取。为应对这些挑战并实现实时重建,我们联合训练姿态网络、深度网络和高斯网络,以预测表示场景的高斯原型。姿态网络与深度网络在无真实深度标签和相机外参条件下,通过自监督方式确定高斯原型的位置。高斯网络则从每张输入图像独立预测原型参数,包括协方差、不透明度及球谐系数。推理阶段,模型可对灵活的多帧环视输入实现前向重建。在nuScenes数据集上的实验表明,本方法在重建质量上优于现有最先进的前馈与场景优化重建方法。

原文摘要 · Abstract (English)

We propose DrivingForward, a feed-forward Gaussian Splatting model that reconstructs driving scenes from flexible surround-view input. Driving scene images from vehicle-mounted cameras are typically sparse, with limited overlap, and the movement of the vehicle further complicates the acquisition of camera extrinsics. To tackle these challenges and achieve real-time reconstruction, we jointly train a pose network, a depth network, and a Gaussian network to predict the Gaussian primitives that represent the driving scenes. The pose network and depth network determine the position of the Gaussian primitives in a self-supervised manner, without using depth ground truth and camera extrinsics during training. The Gaussian network independently predicts primitive parameters from each input image, including covariance, opacity, and spherical harmonics coefficients. At the inference stage, our model can achieve feed-forward reconstruction from flexible multi-frame surround-view input. Experiments on the nuScenes dataset show that our model outperforms existing state-of-the-art feed-forward and scene-optimized reconstruction methods in terms of reconstruction.

3D重建高斯溅射自动驾驶自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。