用动态点图实现视频级4D重建,精度领先。
V-DPM: 4D Video Reconstruction with Dynamic Point Maps
- 将静态点图扩展为视频输入的动态点图,支持神经网络直接预测
- 仅需少量合成数据即可让预训练模型适配动态场景,性能超越现有方法
- 不仅能恢复动态深度,还能捕捉每个点的完整3D运动轨迹
强大的3D表示如DUSt3R的不变点图,能编码3D形状和相机参数,显著推动了前馈式3D重建。尽管点图假设场景静态,动态点图(DPMs)通过额外表示场景运动,拓展了这一概念。然而,现有DPMs仅适用于图像对,且在多视角时需优化后处理,类似DUSt3R。我们主张将DPMs应用于视频,并提出V-DPM验证其有效性。首先,我们设计了一种最大化表征能力、利于神经预测并可复用预训练模型的视频级DPM公式;其次,在VGGT(一种强大的3D重建器)基础上实现该思路。尽管VGGT训练于静态场景,我们仅用少量合成数据即成功将其转化为有效的V-DPM预测器。本方法在动态场景的3D与4D重建上达到当前最优性能。尤其不同于近期的VGGT动态扩展如P3,DPMs不仅能恢复动态深度,还能还原场景中每个点的完整3D运动。
原文摘要 · Abstract (English)
Powerful 3D representations such as DUSt3R invariant point maps, which encode 3D shape and camera parameters, have significantly advanced feed forward 3D reconstruction. While point maps assume static scenes, Dynamic Point Maps (DPMs) extend this concept to dynamic 3D content by additionally representing scene motion. However, existing DPMs are limited to image pairs and, like DUSt3R, require post processing via optimization when more than two views are involved. We argue that DPMs are more useful when applied to videos and introduce V-DPM to demonstrate this. First, we show how to formulate DPMs for video input in a way that maximizes representational power, facilitates neural prediction, and enables reuse of pretrained models. Second, we implement these ideas on top of VGGT, a recent and powerful 3D reconstructor. Although VGGT was trained on static scenes, we show that a modest amount of synthetic data is sufficient to adapt it into an effective V-DPM predictor. Our approach achieves state of the art performance in 3D and 4D reconstruction for dynamic scenes. In particular, unlike recent dynamic extensions of VGGT such as P3, DPMs recover not only dynamic depth but also the full 3D motion of every point in the scene.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。