提出动态点图建模方法,实现自动驾驶中4D场景的高效重建。
DynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous Driving
- 通过共享坐标系联合预测当前与未来点图,隐式学习动态点表示。
- 在真实驾驶数据集上,重建精度显著优于现有方法。
- 适合自动驾驶、动态场景建模等需要实时4D感知的研究者。
自动驾驶中的动态场景重建仍面临巨大挑战,主要源于时间变化剧烈、移动物体多及场景动态复杂。现有前馈3D模型在静态重建中表现良好,但难以捕捉动态运动。为此,我们提出DynamicVGGT,一个统一的前馈框架,将VGGT从静态3D感知扩展至动态4D重建。目标是在前馈3D模型中以动态且时序一致的方式建模点运动。为此,我们在共享参考坐标系中联合预测当前与未来点图,使模型通过时间对应关系隐式学习动态点表示。为高效捕捉时序依赖,引入运动感知时序注意力(MTA)模块,学习运动连续性。此外,设计动态3D高斯泼溅头,通过可学习运动令牌在场景流监督下预测高斯速度,显式建模点运动,并通过连续3D高斯优化精炼动态几何。在自动驾驶数据集上的大量实验表明,DynamicVGGT在复杂驾驶场景下显著优于现有方法,实现了鲁棒的前馈4D动态场景重建。
原文摘要 · Abstract (English)
Dynamic scene reconstruction in autonomous driving remains a fundamental challenge due to significant temporal variations, moving objects, and complex scene dynamics. Existing feed-forward 3D models have demonstrated strong performance in static reconstruction but still struggle to capture dynamic motion. To address these limitations, we propose DynamicVGGT, a unified feed-forward framework that extends VGGT from static 3D perception to dynamic 4D reconstruction. Our goal is to model point motion within feed-forward 3D models in a dynamic and temporally coherent manner. To this end, we jointly predict the current and future point maps within a shared reference coordinate system, allowing the model to implicitly learn dynamic point representations through temporal correspondence. To efficiently capture temporal dependencies, we introduce a Motion-aware Temporal Attention (MTA) module that learns motion continuity. Furthermore, we design a Dynamic 3D Gaussian Splatting Head that explicitly models point motion by predicting Gaussian velocities using learnable motion tokens under scene flow supervision. It refines dynamic geometry through continuous 3D Gaussian optimization. Extensive experiments on autonomous driving datasets demonstrate that DynamicVGGT significantly outperforms existing methods in reconstruction accuracy, achieving robust feed-forward 4D dynamic scene reconstruction under complex driving scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。