arXiv:2510.24734cs.CVcs.LG2025-10被引 2

用两张图实时重建动态驾驶场景,精度效率双提升

DrivingScene: A Multi-Task Online Feed-Forward 3D Gaussian Splatting Method for Dynamic Driving Scenes

  • 基于静态场景先验+轻量残差光流网络,逐相机预测非刚性运动
  • 在nuScenes上实现图像级实时重建,深度、光流、3D点云同步输出
  • 适合自动驾驶感知与实时渲染,尤其擅长复杂动态场景

实时高保真动态驾驶场景重建面临复杂运动与视点稀疏的挑战,现有方法难以兼顾质量与效率。本文提出DrivingScene,一种在线前馈框架,仅需两帧环视图像即可重建4D动态场景。核心创新在于轻量级残差光流网络,在学习到的静态场景先验基础上,对每台相机预测动态物体的非刚性运动,通过场景光流显式建模动态。同时引入粗到精训练范式,避免端到端方法常见的不稳定性。在nuScenes数据集上的实验表明,该图像仅方法可在线生成高质量深度图、场景光流和3D Gaussian点云,显著优于当前最优方法,在动态重建与新视角合成方面表现突出。

原文摘要 · Abstract (English)

Real-time, high-fidelity reconstruction of dynamic driving scenes is challenged by complex dynamics and sparse views, with prior methods struggling to balance quality and efficiency. We propose DrivingScene, an online, feed-forward framework that reconstructs 4D dynamic scenes from only two consecutive surround-view images. Our key innovation is a lightweight residual flow network that predicts the non-rigid motion of dynamic objects per camera on top of a learned static scene prior, explicitly modeling dynamics via scene flow. We also introduce a coarse-to-fine training paradigm that circumvents the instabilities common to end-to-end approaches. Experiments on nuScenes dataset show our image-only method simultaneously generates high-quality depth, scene flow, and 3D Gaussian point clouds online, significantly outperforming state-of-the-art methods in both dynamic reconstruction and novel view synthesis.

3D重建动态场景实时渲染自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。