arXiv:2501.09347cs.CV2025-01被引 2

无需相机姿态信息,用单目视频实现大规模3D重建。

UVRM: A Scalable 3D Reconstruction Model from Unposed Videos

  • 用Transformer隐式聚合视频帧,构建无姿态依赖的特征空间。
  • 在G-Objaverse和CO3D上重建效果接近有姿态标注的方法。
  • 适合缺乏姿态数据的现实场景3D建模,如用户自拍视频重建。

大型重建模型(LRMs)正成为构建3D基础模型的主流方法。传统2D视觉数据训练3D重建模型需依赖已知相机姿态,过程耗时且易出错,导致训练受限于合成3D数据集或带标注姿态的小规模数据集。本文研究了利用无姿态视频进行3D重建的可行性,提出UVRM:一种可在单目视频上训练与评估、无需任何姿态信息的新型3D重建模型。UVRM采用Transformer网络将视频帧隐式聚合至姿态不变的潜在特征空间,并解码为三平面3D表示。为避免训练中依赖真实姿态标注,UVRM结合得分蒸馏采样(SDS)与分析-合成方法,通过预训练扩散模型逐步生成伪新视角。在不依赖姿态信息的前提下,对G-Objaverse和CO3D数据集进行了定性与定量评估。大量实验表明,UVRM能够高效、准确地从无姿态视频中重建多种类3D物体。

原文摘要 · Abstract (English)

Large Reconstruction Models (LRMs) have recently become a popular method for creating 3D foundational models. Training 3D reconstruction models with 2D visual data traditionally requires prior knowledge of camera poses for the training samples, a process that is both time-consuming and prone to errors. Consequently, 3D reconstruction training has been confined to either synthetic 3D datasets or small-scale datasets with annotated poses. In this study, we investigate the feasibility of 3D reconstruction using unposed video data of various objects. We introduce UVRM, a novel 3D reconstruction model capable of being trained and evaluated on monocular videos without requiring any information about the pose. UVRM uses a transformer network to implicitly aggregate video frames into a pose-invariant latent feature space, which is then decoded into a tri-plane 3D representation. To obviate the need for ground-truth pose annotations during training, UVRM employs a combination of the score distillation sampling (SDS) method and an analysis-by-synthesis approach, progressively synthesizing pseudo novel-views using a pre-trained diffusion model. We qualitatively and quantitatively evaluate UVRM's performance on the G-Objaverse and CO3D datasets without relying on pose information. Extensive experiments show that UVRM is capable of effectively and efficiently reconstructing a wide range of 3D objects from unposed videos.

3D重建单目视频无姿态扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。