arXiv:2509.13652cs.CV2025-09中稿 · AJCAI 2025

用单视图重建直接对齐3D场景,实现精确相机位姿估计

Gaussian Alignment for Relative Camera Pose Estimation via Single-View Reconstruction

  • 通过独立重建两幅图像的3D高斯混合模型并直接对齐
  • 在RealEstate10K上超越MASt3R等主流方法,误差降低23%
  • 无需特征匹配,适合纹理缺失或大视差场景

从一对图像中估计度量相对相机位姿在三维重建与定位中至关重要。然而,传统两视图位姿估计方法仅能获得尺度不确定的位姿,且在大基线、无纹理或反光表面下表现不佳。本文提出GARPS——一种无需训练的框架,将该问题转化为两个独立重建的3D场景之间的直接对齐。GARPS利用度量单目深度估计器和高斯场景重建器,为每张图像生成度量3D高斯混合模型(GMM),再通过优化可微的GMM对齐目标,精炼来自前馈式两视图位姿估计器的初始位姿。该目标联合考虑几何结构、视角无关颜色、各向异性协方差及语义特征一致性,对遮挡和纹理贫乏区域具有鲁棒性,且无需显式2D对应关系。在RealEstate10K数据集上的大量实验表明,GARPS优于经典及最新学习型方法,包括MASt3R。结果凸显了将单视图感知与多视图几何结合以实现鲁棒、度量相对位姿估计的潜力。

原文摘要 · Abstract (English)

Estimating metric relative camera pose from a pair of images is of great importance for 3D reconstruction and localisation. However, conventional two-view pose estimation methods are not metric, with camera translation known only up to a scale, and struggle with wide baselines and textureless or reflective surfaces. This paper introduces GARPS, a training-free framework that casts this problem as the direct alignment of two independently reconstructed 3D scenes. GARPS leverages a metric monocular depth estimator and a Gaussian scene reconstructor to obtain a metric 3D Gaussian Mixture Model (GMM) for each image. It then refines an initial pose from a feed-forward two-view pose estimator by optimising a differentiable GMM alignment objective. This objective jointly considers geometric structure, view-independent colour, anisotropic covariance, and semantic feature consistency, and is robust to occlusions and texture-poor regions without requiring explicit 2D correspondences. Extensive experiments on the Real\-Estate10K dataset demonstrate that GARPS outperforms both classical and state-of-the-art learning-based methods, including MASt3R. These results highlight the potential of bridging single-view perception with multi-view geometry to achieve robust and metric relative pose estimation.

位姿估计3D重建高斯表示单视图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。