用视频生成模型做单目动态场景三维重建,零样本泛化效果好。
Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction

- 利用预训练视频模型的动态先验,仅用合成数据训练
- 预测点、视差、光线三类几何图,精度显著优于现有方法
- 适合需要快速部署、无真实标注数据的动态场景重建任务
我们提出Geo4D,一种将视频扩散模型用于单目动态场景三维重建的方法。通过利用大规模预训练视频模型所捕捉的强大动态先验,Geo4D仅需合成数据即可训练,并在真实数据上实现零样本泛化。该方法预测点图、视差图和光线图等多种互补几何模态,提出新的多模态对齐算法以融合这些信息,并在推理时采用滑动窗口策略,从而实现对长视频的鲁棒且精确的4D重建。在多个基准测试中,Geo4D显著超越当前最优视频深度估计方法。
原文摘要 · Abstract (English)
We introduce Geo4D, a method to repurpose video diffusion models for monocular 3D reconstruction of dynamic scenes. By leveraging the strong dynamic priors captured by large-scale pre-trained video models, Geo4D can be trained using only synthetic data while generalizing well to real data in a zero-shot manner. Geo4D predicts several complementary geometric modalities, namely point, disparity, and ray maps. We propose a new multi-modal alignment algorithm to align and fuse these modalities, as well as a sliding window approach at inference time, thus enabling robust and accurate 4D reconstruction of long videos. Extensive experiments across multiple benchmarks show that Geo4D significantly surpasses state-of-the-art video depth estimation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。