用百万级360视频训练模型,实现真实场景的自由视角生成。
From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos
- 基于360-1M数据集,通过扩散模型学习多视角对应帧。
- 在真实场景中实现自由相机移动与新视角生成,精度优于基准方法。
- 适合做3D理解、场景重建和虚拟现实的研究者使用。
三维(3D)物体与场景理解是人类与世界交互能力的核心,在计算机视觉、图形学和机器人领域备受关注。大规模合成与以物体为中心的3D数据集已被证明能有效训练具备3D理解能力的模型。然而,将类似方法应用于真实世界对象与场景时面临数据规模不足的挑战。视频可作为真实世界3D数据的潜在来源,但大规模获取同一内容的多样化且对应的视角仍具难度。此外,标准视频具有固定视角,限制了从多种多样视角访问场景的能力。本文提出,大规模360视频可克服上述局限,提供可扩展的多视角对应帧。我们构建了360-1M数据集,并设计了高效获取跨视角对应帧的流程。在此基础上,我们训练了基于扩散模型的Odin模型。依托迄今最大的真实世界多视角数据集,Odin能够自由生成真实场景的新视角。不同于以往方法,Odin可实现相机在环境中的移动,使模型能推断场景的几何结构与布局。此外,我们在标准的新视角合成与3D重建基准测试中均取得性能提升。
原文摘要 · Abstract (English)
Three-dimensional (3D) understanding of objects and scenes play a key role in humans' ability to interact with the world and has been an active area of research in computer vision, graphics, and robotics. Large scale synthetic and object-centric 3D datasets have shown to be effective in training models that have 3D understanding of objects. However, applying a similar approach to real-world objects and scenes is difficult due to a lack of large-scale data. Videos are a potential source for real-world 3D data, but finding diverse yet corresponding views of the same content has shown to be difficult at scale. Furthermore, standard videos come with fixed viewpoints, determined at the time of capture. This restricts the ability to access scenes from a variety of more diverse and potentially useful perspectives. We argue that large scale 360 videos can address these limitations to provide: scalable corresponding frames from diverse views. In this paper, we introduce 360-1M, a 360 video dataset, and a process for efficiently finding corresponding frames from diverse viewpoints at scale. We train our diffusion-based model, Odin, on 360-1M. Empowered by the largest real-world, multi-view dataset to date, Odin is able to freely generate novel views of real-world scenes. Unlike previous methods, Odin can move the camera through the environment, enabling the model to infer the geometry and layout of the scene. Additionally, we show improved performance on standard novel view synthesis and 3D reconstruction benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。