用360°视频生成可行走的3D场景,解决传统方法视野受限问题。
Generating 360° Video is What You Need For a 3D Scene
- 以360°视频为中间表示,实现完整场景建模
- 128帧视频模拟行走路径,重建后导航流畅度提升
- 适合需要高沉浸感3D内容的开发者与创作者
由于缺乏现成的场景数据,3D场景生成仍是难题。现有方法多生成局部场景,导航自由度有限。本文提出一种实用且可扩展的方案:以360°视频作为中间场景表示,捕捉完整场景上下文并保证视觉一致性。我们设计了WorldPrompter,一个从文本提示生成可漫游3D场景的生成流程。该流程包含一个条件生成的360°全景视频生成器,可生成128帧视频,模拟人行进中拍摄虚拟环境。随后通过快速前馈式3D重建器将视频转为高斯点云,实现真实可行走体验。实验表明,该模型在图像与视频混合数据上训练,对静态场景实现良好时空一致性,平均COLMAP匹配率达94.6%,支持高质量全景高斯点云重建,显著提升场景内导航效果。定性与定量结果均显示其优于当前最先进的360°视频生成与3D场景生成模型。
原文摘要 · Abstract (English)
Generating 3D scenes is still a challenging task due to the lack of readily available scene data. Most existing methods only produce partial scenes and provide limited navigational freedom. We introduce a practical and scalable solution that uses 360° video as an intermediate scene representation, capturing the full-scene context and ensuring consistent visual content throughout the generation. We propose WorldPrompter, a generative pipeline that synthesizes traversable 3D scenes from text prompts. WorldPrompter incorporates a conditional 360° panoramic video generator, capable of producing a 128-frame video that simulates a person walking through and capturing a virtual environment. The resulting video is then reconstructed as Gaussian splats by a fast feedforward 3D reconstructor, enabling a true walkable experience within the 3D scene. Experiments demonstrate that our panoramic video generation model, trained with a mix of image and video data, achieves convincing spatial and temporal consistency for static scenes. This is validated by an average COLMAP matching rate of 94.6\%, allowing for high-quality panoramic Gaussian splat reconstruction and improved navigation throughout the scene. Qualitative and quantitative results also show it outperforms the state-of-the-art 360° video generators and 3D scene generation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。