用视频生成模型直接构建可扩展的3D场景,避免几何失真。
WonderVerse: Extendable 3D Scene Generation with Video Generative Models
- 利用视频生成模型的全局先验生成3D场景
- 支持可控扩展,生成更大更连贯的环境
- 适合需要高质量3D内容的创作者和开发者
我们提出WonderVerse,一个简单而高效的可扩展3D场景生成框架。与依赖迭代深度估计和图像修复、易产生几何畸变和不一致性的现有方法不同,WonderVerse借助视频生成基础模型中嵌入的世界级先验,生成高度沉浸且几何一致的3D环境。此外,我们提出一种新的可控3D场景扩展技术,显著提升生成环境的规模;引入一种基于相机轨迹的异常序列检测模块,解决生成视频中的几何不一致性问题。最后,WonderVerse兼容多种3D重建方法,实现高效且高质量的生成。大量实验表明,凭借简洁的流程,WonderVerse生成的可扩展、高保真3D场景,明显优于依赖复杂架构的现有工作。
原文摘要 · Abstract (English)
We introduce \textit{WonderVerse}, a simple but effective framework for generating extendable 3D scenes. Unlike existing methods that rely on iterative depth estimation and image inpainting, often leading to geometric distortions and inconsistencies, WonderVerse leverages the powerful world-level priors embedded within video generative foundation models to create highly immersive and geometrically coherent 3D environments. Furthermore, we propose a new technique for controllable 3D scene extension to substantially increase the scale of the generated environments. Besides, we introduce a novel abnormal sequence detection module that utilizes camera trajectory to address geometric inconsistency in the generated videos. Finally, WonderVerse is compatible with various 3D reconstruction methods, allowing both efficient and high-quality generation. Extensive experiments on 3D scene generation demonstrate that our WonderVerse, with an elegant and simple pipeline, delivers extendable and highly-realistic 3D scenes, markedly outperforming existing works that rely on more complex architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。