arXiv:2510.10097cs.CV2025-10

无需精确相机位姿,用图像几何信息实现稀疏视角下稳定3D重建

Gesplat: Robust Pose-Free 3D Reconstruction via Geometry-Guided Gaussian Splatting

  • 用VGGT模型替代COLMAP生成更可靠的初始姿态和稠密点云
  • 通过视图间一致性优化与流式深度正则化,提升重建精度与鲁棒性
  • 适合缺乏精确位姿信息的复杂场景或稀疏拍摄条件下的3D建模

神经辐射场(NeRF)和3D高斯溅射(3DGS)虽推动了3D重建与新视角合成的发展,但严重依赖准确的相机位姿和密集视角覆盖,限制了其在稀疏视角场景中的应用。为此,我们提出Gesplat,一种基于3DGS的框架,可在无位姿约束的稀疏图像下实现鲁棒的新视角合成与几何一致的重建。不同于以往依赖COLMAP进行稀疏点云初始化的方法,我们采用VGGT基础模型获取更可靠的初始位姿与稠密点云。本方法包含三项关键创新:1)结合双位置-形状优化的混合高斯表示,辅以视图间匹配一致性;2)基于图结构的属性精修模块,增强场景细节;3)基于流的深度正则化,提升深度估计精度以实现更有效的监督。大量定量与定性实验表明,相比其他无位姿方法,Gesplat在前向面对与大规模复杂数据集上均表现更优。

原文摘要 · Abstract (English)

Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have advanced 3D reconstruction and novel view synthesis, but remain heavily dependent on accurate camera poses and dense viewpoint coverage. These requirements limit their applicability in sparse-view settings, where pose estimation becomes unreliable and supervision is insufficient. To overcome these challenges, we introduce Gesplat, a 3DGS-based framework that enables robust novel view synthesis and geometrically consistent reconstruction from unposed sparse images. Unlike prior works that rely on COLMAP for sparse point cloud initialization, we leverage the VGGT foundation model to obtain more reliable initial poses and dense point clouds. Our approach integrates several key innovations: 1) a hybrid Gaussian representation with dual position-shape optimization enhanced by inter-view matching consistency; 2) a graph-guided attribute refinement module to enhance scene details; and 3) flow-based depth regularization that improves depth estimation accuracy for more effective supervision. Comprehensive quantitative and qualitative experiments demonstrate that our approach achieves more robust performance on both forward-facing and large-scale complex datasets compared to other pose-free methods.

3D重建高斯溅射无位姿几何引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。