arXiv:2602.18906cs.CV2026-02被引 3

用密集深度图提升单目相机位姿估计精度,效果超越当前最优

Marginalized Bundle Adjustment: Multi-View Camera Pose from Monocular Depth Estimates

  • 基于深度图密度设计新型优化方法,降低误差影响
  • 在多视角场景中实现媲美或超过现有最佳的定位性能
  • 适用于从少帧到上千张图像的各类复杂场景

结构光重建(SfM)是从多视角图像中恢复相机参数与场景几何的基础任务。尽管深度学习使单图像单目深度估计(MDE)无需依赖相机运动即可实现高精度,但将MDE融入SfM仍面临挑战。与传统三角化生成的稀疏点云不同,MDE输出稠密深度图,误差方差显著更高。受现代RANSAC估计算法启发,我们提出边缘化束调整(MBA),利用深度图密度特性缓解误差方差。实验表明,借助MBA,MDE深度图足以在SfM和相机重定位任务中达到或超越当前最优结果。在从少量图像到包含数千张图像的大规模系统上,均展现出稳定鲁棒的表现。本方法凸显了MDE在多视角3D视觉中的巨大潜力。

原文摘要 · Abstract (English)

Structure-from-Motion (SfM) is a fundamental 3D vision task for recovering camera parameters and scene geometry from multi-view images. While recent deep learning advances enable accurate Monocular Depth Estimation (MDE) from single images without depending on camera motion, integrating MDE into SfM remains a challenge. Unlike conventional triangulated sparse point clouds, MDE produces dense depth maps with significantly higher error variance. Inspired by modern RANSAC estimators, we propose Marginalized Bundle Adjustment (MBA) to mitigate MDE error variance leveraging its density. With MBA, we show that MDE depth maps are sufficiently accurate to yield SoTA or competitive results in SfM and camera relocalization tasks. Through extensive evaluations, we demonstrate consistently robust performance across varying scales, ranging from few-frame setups to large multi-view systems with thousands of images. Our method highlights the significant potential of MDE in multi-view 3D vision.

三维重建深度估计位姿估计束调整

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。