arXiv:2412.12091cs.CV2024-12CVPR被引 108

仅用一张图高效生成高质量3D场景,突破传统方法局限。

Wonderland: Navigating 3D Scenes from a Single Image

论文配图:Wonderland: Navigating 3D Scenes from a Single Image
图 1 · 摘自论文原文
  • 利用视频扩散模型的潜空间实现前馈式3D重建
  • 在多数据集上显著优于现有单图3D生成方法
  • 适合需要快速生成通用3D场景的研究与应用

如何从任意单张图像高效生成高质量、大范围3D场景?现有方法存在需多视角数据、每场景优化耗时、遮挡区域几何失真及背景视觉质量低等缺陷。本文提出新型3D场景重建流程,引入大规模重建模型,利用视频扩散模型的潜变量预测3D高斯泼溅(3D Gaussian Splattings)。该视频扩散模型专为精确遵循指定相机轨迹生成视频而设计,可生成压缩视频潜变量,编码多视角信息并保持3D一致性。通过渐进式学习策略训练3D重建模型在视频潜空间操作,实现高效生成高质量、大范围、通用性强的3D场景。跨多个数据集的广泛评估表明,本模型显著优于现有单图3D场景生成方法,尤其在域外图像上表现突出。首次证明可基于扩散模型潜空间构建3D重建模型,实现高效3D场景生成。

原文摘要 · Abstract (English)

How can one efficiently generate high-quality, wide-scope 3D scenes from arbitrary single images? Existing methods suffer several drawbacks, such as requiring multi-view data, time-consuming per-scene optimization, distorted geometry in occluded areas, and low visual quality in backgrounds. Our novel 3D scene reconstruction pipeline overcomes these limitations to tackle the aforesaid challenge. Specifically, we introduce a large-scale reconstruction model that leverages latents from a video diffusion model to predict 3D Gaussian Splattings of scenes in a feed-forward manner. The video diffusion model is designed to create videos precisely following specified camera trajectories, allowing it to generate compressed video latents that encode multi-view information while maintaining 3D consistency. We train the 3D reconstruction model to operate on the video latent space with a progressive learning strategy, enabling the efficient generation of high-quality, wide-scope, and generic 3D scenes. Extensive evaluations across various datasets affirm that our model significantly outperforms existing single-view 3D scene generation methods, especially with out-of-domain images. Thus, we demonstrate for the first time that a 3D reconstruction model can effectively be built upon the latent space of a diffusion model in order to realize efficient 3D scene generation.

3D重建单图生成扩散模型高斯泼溅

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。