融合不确定性建模的单目位姿估计算法,提升定位精度与稳定性。
VKFPos: A Learning-Based Monocular Positioning with Variational Bayesian Extended Kalman Filter Integration
- 通过变分贝叶斯框架整合绝对与相对位姿回归
- 单帧定位达顶尖水平,时序定位显著优于现有方法
- 适合对精度和鲁棒性要求高的视觉定位场景
本文针对基于学习的单目位姿估计挑战,提出VKFPos,一种将绝对位姿回归(APR)与相对位姿回归(RPR)通过扩展卡尔曼滤波(EKF)嵌入变分贝叶斯推断框架的新方法。研究发现,单目定位问题的后验概率可分解为APR与RPR两部分,该分解通过在APR和RPR分支中预测协方差实现,使模型能显式建模不确定性,并用于优化损失函数与EKF融合。在室内与室外数据集上的实验表明,单次推理的APR分支性能达到当前最优水平;而在时序定位任务中,结合连续图像的RPR与EKF后,VKFPos显著优于时序APR及传统模型融合方法,定位精度更优。
原文摘要 · Abstract (English)
This paper addresses the challenges in learning-based monocular positioning by proposing VKFPos, a novel approach that integrates Absolute Pose Regression (APR) and Relative Pose Regression (RPR) via an Extended Kalman Filter (EKF) within a variational Bayesian inference framework. Our method shows that the essential posterior probability of the monocular positioning problem can be decomposed into APR and RPR components. This decomposition is embedded in the deep learning model by predicting covariances in both APR and RPR branches, allowing them to account for associated uncertainties. These covariances enhance the loss functions and facilitate EKF integration. Experimental evaluations on both indoor and outdoor datasets show that the single-shot APR branch achieves accuracy on par with state-of-the-art methods. Furthermore, for temporal positioning, where consecutive images allow for RPR and EKF integration, VKFPos outperforms temporal APR and model-based integration methods, achieving superior accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。