arXiv:2411.13549cs.CV2024-11CVPR被引 7

用无姿态网络照片生成3D一致视频,无需3D标注

KFC-W: Generating 3D-Consistent Videos from Unposed Internet Photos

  • 自监督学习利用多视角照片的多样性与视频一致性
  • 在几何与外观一致性上超越所有基线模型
  • 适合需要相机控制的3D重建应用

我们解决从无姿态网络照片生成视频的问题。少量输入图像作为关键帧,模型在其间插值以模拟摄像机移动路径。给定随机图像,模型捕捉潜在几何结构、识别场景身份并关联帧间的相机位置与方向,体现了对3D结构和场景布局的根本理解。然而,现有视频模型如Luma Dream Machine在此任务上表现不佳。我们设计了一种自监督方法,利用视频的一致性与多视角网络照片的变异性,在无需任何3D标注(如相机参数)的情况下训练可扩展的3D感知视频模型。验证表明,该方法在几何与外观一致性方面优于所有基线。此外,我们的模型显著提升需相机控制的应用(如3D Gaussian Splatting)性能。结果表明,仅使用2D数据(如视频和多视角网络照片)即可实现场景级3D学习的规模化。

原文摘要 · Abstract (English)

We address the problem of generating videos from unposed internet photos. A handful of input images serve as keyframes, and our model interpolates between them to simulate a path moving between the cameras. Given random images, a model's ability to capture underlying geometry, recognize scene identity, and relate frames in terms of camera position and orientation reflects a fundamental understanding of 3D structure and scene layout. However, existing video models such as Luma Dream Machine fail at this task. We design a self-supervised method that takes advantage of the consistency of videos and variability of multiview internet photos to train a scalable, 3D-aware video model without any 3D annotations such as camera parameters. We validate that our method outperforms all baselines in terms of geometric and appearance consistency. We also show our model benefits applications that enable camera control, such as 3D Gaussian Splatting. Our results suggest that we can scale up scene-level 3D learning using only 2D data such as videos and multiview internet photos.

视频生成3D一致性自监督学习多视角重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。