arXiv:2512.03004cs.CV2025-12被引 22

无需相机位姿即可快速重建动态驾驶场景,支持长序列多视角输入。

DGGT: Feedforward 4D Reconstruction of Dynamic Driving Scenes using Unposed Images

  • 将相机位姿作为模型输出而非输入,实现端到端无标定重建。
  • 在Waymo、nuScenes等数据集上达到最优性能,且输入帧数增加时仍高效扩展。
  • 适合自动驾驶系统训练与评估,尤其适用于稀疏、未标定的视频输入。

自动驾驶需要快速、可扩展的4D场景重建与重仿真以用于训练和评估,但现有动态驾驶场景方法仍依赖逐场景优化、已知相机标定或短时序窗口,导致速度慢且不实用。本文从前馈视角重新审视该问题,提出 extbf{Driving Gaussian Grounded Transformer (DGGT)},一种统一的无位姿动态场景重建框架。我们发现,将相机位姿视为必要输入会限制灵活性与可扩展性。因此,我们将其重构为模型输出,使重建可直接从稀疏、未标定图像出发,并支持任意数量视角和长序列输入。本方法联合预测每帧3D高斯图与相机参数,通过轻量级动态头解耦运动信息,利用寿命头在时间维度上调节可见性以保持一致性;扩散渲染精修进一步降低运动/插值伪影,提升稀疏输入下的新视角质量。最终实现单次前馈、无位姿的算法,在大规模驾驶基准(Waymo、nuScenes、Argoverse2)上均超越现有方法,无论在各数据集上训练还是跨数据集零样本迁移,表现均领先,且随输入帧数增加仍具良好扩展性。

原文摘要 · Abstract (English)

Autonomous driving needs fast, scalable 4D reconstruction and re-simulation for training and evaluation, yet most methods for dynamic driving scenes still rely on per-scene optimization, known camera calibration, or short frame windows, making them slow and impractical. We revisit this problem from a feedforward perspective and introduce \textbf{Driving Gaussian Grounded Transformer (DGGT)}, a unified framework for pose-free dynamic scene reconstruction. We note that the existing formulations, treating camera pose as a required input, limit flexibility and scalability. Instead, we reformulate pose as an output of the model, enabling reconstruction directly from sparse, unposed images and supporting an arbitrary number of views for long sequences. Our approach jointly predicts per-frame 3D Gaussian maps and camera parameters, disentangles dynamics with a lightweight dynamic head, and preserves temporal consistency with a lifespan head that modulates visibility over time. A diffusion-based rendering refinement further reduces motion/interpolation artifacts and improves novel-view quality under sparse inputs. The result is a single-pass, pose-free algorithm that achieves state-of-the-art performance and speed. Trained and evaluated on large-scale driving benchmarks (Waymo, nuScenes, Argoverse2), our method outperforms prior work both when trained on each dataset and in zero-shot transfer across datasets, and it scales well as the number of input frames increases.

4D重建自动驾驶无标定高斯溅射

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。