用全帧几何特征提升3D场景生成质量与速度
GeoWorld: Providing Full-frame Geometry Features to Facilitate 3D Scene Generation

- 分两阶段生成:先建几何框架,再基于框架生成3D内容
- 生成的3D场景更清晰,比当前最佳方法快7.5倍
- 适合需要高质量3D生成的视觉应用开发者
以往利用视频模型进行图像到3D场景生成的方法常出现几何失真和模糊问题。仅凭单帧输入依赖视频模型隐式保持几何一致性效果不佳。本文提出两阶段方法GeoWorld,通过第一阶段视频生成模型结合多视图几何模型,输出全帧几何特征,作为第二阶段生成模型的几何参考。引入几何损失以施加真实世界几何约束,并设计几何适配模块确保特征有效利用。得益于全帧几何建模,本方法使用两个较小的视频模型,在生成更高保真度3D场景的同时,速度优于现有最佳方法,例如比Hunyuan-Voyager快7.5倍。
原文摘要 · Abstract (English)
Previous works that leverage video models for image-to-3D scene generation often suffer from geometric distortions and blurry content. Using video generation models to implicitly maintain geometric consistency according to a single-frame input is ineffective. In this paper, we present a two-stage method, named $\textbf{GeoWorld}$, that renovates the image-to-3D scene generation pipeline by providing full-frame geometry features. The first-stage video generation model, followed by a multi-view geometry model, produces $\textbf{full-frame}$ geometry features, which are then used as a mental draft of geometric conditions to aid the second-stage video-generation model. A geometric loss is proposed to impose real-world geometric constraints, and a geometry adaptation module is introduced to ensure the effective utilization of geometry features. Thanks to full-frame geometric modeling, the two smaller video models in our two-stage method can generate higher-fidelity 3D scenes than SOTA methods, while being even faster, e.g. 7.5$\times$ faster than Hunyuan-Voyager. Project page: https://peaes.github.io/GeoWorld.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。