仅用一张图重建沉浸式3D场景,突破视角与质量限制。
ExScene: Free-View 3D Scene Reconstruction with Gaussian Splatting from a Single Image
- 分两阶段:先生成全景图再构建3D高斯点云
- 结合扩散模型与几何信息,实现高保真全景重建
- 适合虚拟现实、游戏建模等需要高质量3D内容的场景
增强现实与虚拟现实应用对从单张图像生成沉浸式3D场景的需求日益增长。然而,由于单视图输入提供的先验信息有限,现有方法通常只能重建一致性差、视野狭窄的3D场景,难以泛化至沉浸式场景重建。为此,我们提出ExScene,一种两阶段流程,可从任意单视图图像重建沉浸式3D场景。ExScene设计了一种新型多模态扩散模型,生成高保真且全局一致的全景图像;随后提出全景深度估计方法,从全景图中计算几何信息,并结合高保真全景图训练初始3D高斯点云(3DGS)模型。接着引入基于2D稳定视频扩散先验的GS精炼技术,在去噪过程中加入相机轨迹一致性与色彩-几何先验,提升图像序列的颜色与空间一致性。最终使用优化后的序列微调初始3DGS模型,显著提升重建质量。实验表明,ExScene仅依赖单视图输入即可实现一致且沉浸式的场景重建,显著优于当前最先进方法。
原文摘要 · Abstract (English)
The increasing demand for augmented and virtual reality applications has highlighted the importance of crafting immersive 3D scenes from a simple single-view image. However, due to the partial priors provided by single-view input, existing methods are often limited to reconstruct low-consistency 3D scenes with narrow fields of view from single-view input. These limitations make them less capable of generalizing to reconstruct immersive scenes. To address this problem, we propose ExScene, a two-stage pipeline to reconstruct an immersive 3D scene from any given single-view image. ExScene designs a novel multimodal diffusion model to generate a high-fidelity and globally consistent panoramic image. We then develop a panoramic depth estimation approach to calculate geometric information from panorama, and we combine geometric information with high-fidelity panoramic image to train an initial 3D Gaussian Splatting (3DGS) model. Following this, we introduce a GS refinement technique with 2D stable video diffusion priors. We add camera trajectory consistency and color-geometric priors into the denoising process of diffusion to improve color and spatial consistency across image sequences. These refined sequences are then used to fine-tune the initial 3DGS model, leading to better reconstruction quality. Experimental results demonstrate that our ExScene achieves consistent and immersive scene reconstruction using only single-view input, significantly surpassing state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。