无需相机位姿,用扩散模型重建360度场景,效果媲美有位姿的先进方法。
Gaussian Scenes: Pose-Free Sparse-View Scene Reconstruction using Depth-Enhanced Diffusion Priors
- 用图像到图像扩散模型修复新视角渲染和深度图中的缺失细节与伪影
- 在MipNeRF360和DL3DV-10K上优于现有无位姿方法,接近有位姿最优水平
- 引入轻量级几何与上下文条件控制,适合复杂场景的无位姿重建
本文提出一种生成式方法,实现从少量2D图像中进行无相机位姿的360度场景重建。传统方法依赖深度估计或3D先验来约束不完整、无位姿的观测,但现有基于视图条件的生成先验方法需已知相机位姿,无法直接用于无位姿场景。为此,我们设计了一种图像到图像生成模型,用于修复3D场景的新视角渲染和深度图中的缺失内容与伪影。通过特征逐元素线性调制(FiLM)实现上下文与几何条件控制,替代复杂的交叉注意力机制;并提出针对3D高斯点云表示的新置信度度量,提升伪影检测能力。采用受高斯-SLAM启发的渐进式融合策略,构建多视角一致的3D表示。在MipNeRF360与DL3DV-10K基准测试中,本方法超越现有无位姿方法,且在复杂360场景中表现接近具有预计算位姿的最先进方法。
原文摘要 · Abstract (English)
In this work, we introduce a generative approach for pose-free (without camera parameters) reconstruction of 360 scenes from a sparse set of 2D images. Pose-free scene reconstruction from incomplete, pose-free observations is usually regularized with depth estimation or 3D foundational priors. While recent advances have enabled sparse-view reconstruction of large complex scenes (with high degree of foreground and background detail) with known camera poses using view-conditioned generative priors, these methods cannot be directly adapted for the pose-free setting when ground-truth poses are not available during evaluation. To address this, we propose an image-to-image generative model designed to inpaint missing details and remove artifacts in novel view renders and depth maps of a 3D scene. We introduce context and geometry conditioning using Feature-wise Linear Modulation (FiLM) modulation layers as a lightweight alternative to cross-attention and also propose a novel confidence measure for 3D Gaussian splat representations to allow for better detection of these artifacts. By progressively integrating these novel views in a Gaussian-SLAM-inspired process, we achieve a multi-view-consistent 3D representation. Evaluations on the MipNeRF360 and DL3DV-10K benchmark dataset demonstrate that our method surpasses existing pose-free techniques and performs competitively with state-of-the-art posed (precomputed camera parameters are given) reconstruction methods in complex 360 scenes. Our project page provides additional results, videos, and code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。