arXiv:2501.05446cs.CV2025-01CVPR被引 17

用仿射校正提升单目深度先验的相对位姿估计性能

Relative Pose Estimation through Affine Corrections of Monocular Depth Priors

  • 显式建模单目深度的尺度与偏移不确定性,适配校准与非校准场景
  • 在多个数据集上显著超越传统关键点与PnP方法,提升幅度达15%以上
  • 适用于各类特征匹配器与深度模型,对最新进展也有良好兼容性

近年来单目深度估计(MDE)模型取得显著进展,许多模型可预测仿射不变的相对深度,而大规模训练与视觉基础模型使合理估计度量(绝对)深度成为可能。然而,如何有效利用这些预测进行几何视觉任务,特别是相对位姿估计,仍研究不足。尽管深度提供丰富的跨视图对齐约束,但单目深度先验的固有噪声与模糊性给改进经典基于关键点的方法带来实际挑战。本文提出三种显式处理独立仿射(尺度与偏移)模糊性的相对位姿估计求解器,覆盖校准与非校准条件。进一步设计一种融合所提求解器与经典点匹配求解器及对极约束的混合流程。实验表明,仿射校正不仅有助于相对深度先验,甚至对“度量”深度也带来意外提升。在多个数据集上的结果均显示,本方法在校准与非校准设置下显著优于经典关键点基线与基于PnP的方案。此外,该方法在不同特征匹配器与MDE模型下均保持稳定提升,并能受益于两项模块的最新进展。代码已公开于https://github.com/MarkYu98/madpose。

原文摘要 · Abstract (English)

Monocular depth estimation (MDE) models have undergone significant advancements over recent years. Many MDE models aim to predict affine-invariant relative depth from monocular images, while recent developments in large-scale training and vision foundation models enable reasonable estimation of metric (absolute) depth. However, effectively leveraging these predictions for geometric vision tasks, in particular relative pose estimation, remains relatively under explored. While depths provide rich constraints for cross-view image alignment, the intrinsic noise and ambiguity from the monocular depth priors present practical challenges to improving upon classic keypoint-based solutions. In this paper, we develop three solvers for relative pose estimation that explicitly account for independent affine (scale and shift) ambiguities, covering both calibrated and uncalibrated conditions. We further propose a hybrid estimation pipeline that combines our proposed solvers with classic point-based solvers and epipolar constraints. We find that the affine correction modeling is beneficial to not only the relative depth priors but also, surprisingly, the "metric" ones. Results across multiple datasets demonstrate large improvements of our approach over classic keypoint-based baselines and PnP-based solutions, under both calibrated and uncalibrated setups. We also show that our method improves consistently with different feature matchers and MDE models, and can further benefit from very recent advances on both modules. Code is available at https://github.com/MarkYu98/madpose.

位姿估计单目深度仿射校正几何视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。