arXiv:2503.21082cs.CV2025-03被引 9

用视频扩散模型直接重建动态3D场景,无需额外模块。

Can Video Diffusion Model Reconstruct 4D Geometry?

  • 利用预训练视频扩散模型的时空先验,直接生成4D点云图。
  • 在多种动态场景下性能媲美顶尖方法,可准确恢复相机位姿与几何细节。
  • 全流程前馈,不依赖深度、光流等外部模块,适合快速重建任务。

从单目视频重建动态3D场景(即4D几何)是一个重要但具挑战性的问题。传统多视角几何方法在动态运动下表现不佳,而近年学习型方法或需专用4D表示,或依赖复杂优化。本文提出Sora3R框架,利用大规模视频扩散模型丰富的时空先验,直接从普通视频中推断4D点云图。Sora3R采用两阶段流程:(1) 基于预训练视频VAE适配点云图VAE,确保几何与视频隐空间兼容;(2) 在视频与点云图隐空间联合微调扩散主干,为每帧生成一致的4D点云图。该方法完全前馈运行,无需外部模块(如深度、光流或分割)或迭代全局对齐。大量实验表明,Sora3R能可靠恢复相机位姿与详细场景几何,在多样动态场景中性能达到当前最佳水平。

原文摘要 · Abstract (English)

Reconstructing dynamic 3D scenes (i.e., 4D geometry) from monocular video is an important yet challenging problem. Conventional multiview geometry-based approaches often struggle with dynamic motion, whereas recent learning-based methods either require specialized 4D representation or sophisticated optimization. In this paper, we present Sora3R, a novel framework that taps into the rich spatiotemporal priors of large-scale video diffusion models to directly infer 4D pointmaps from casual videos. Sora3R follows a two-stage pipeline: (1) we adapt a pointmap VAE from a pretrained video VAE, ensuring compatibility between the geometry and video latent spaces; (2) we finetune a diffusion backbone in combined video and pointmap latent space to generate coherent 4D pointmaps for every frame. Sora3R operates in a fully feedforward manner, requiring no external modules (e.g., depth, optical flow, or segmentation) or iterative global alignment. Extensive experiments demonstrate that Sora3R reliably recovers both camera poses and detailed scene geometry, achieving performance on par with state-of-the-art methods for dynamic 4D reconstruction across diverse scenarios.

4D重建扩散模型视频生成点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。