用深度不确定性提升室内定位精度,无需为每个场景重训练模型。
UnLoc: Leveraging Depth Uncertainties for Floorplan Localization
- 将深度预测建模为概率分布,显式捕捉不确定性。
- 在长序列上定位召回率提升2.7倍,短序列达42.2倍。
- 可直接使用预训练单目深度模型,通用性强,适合部署于新环境。
我们提出UnLoc,一种高效的基于数据驱动的楼层平面内连续相机定位方法。楼层平面数据易于获取、长期稳定且对视觉变化鲁棒。针对现有方法缺乏深度预测不确定性建模及需为每环境定制深度网络的问题,我们引入一种新型概率模型,将深度预测表示为显式概率分布。通过利用现成的预训练单目深度模型,避免了对每环境训练专用深度网络的需求,显著提升对未见空间的泛化能力。我们在大规模合成与真实世界数据集上评估UnLoc,结果表明其在准确性和鲁棒性方面均优于现有方法。尤其在具有挑战性的LaMAR HGE数据集上,长序列(100帧)定位召回率比最先进方法高2.7倍,短序列(15帧)高达42.2倍。
原文摘要 · Abstract (English)
We propose UnLoc, an efficient data-driven solution for sequential camera localization within floorplans. Floorplan data is readily available, long-term persistent, and robust to changes in visual appearance. We address key limitations of recent methods, such as the lack of uncertainty modeling in depth predictions and the necessity for custom depth networks trained for each environment. We introduce a novel probabilistic model that incorporates uncertainty estimation, modeling depth predictions as explicit probability distributions. By leveraging off-the-shelf pre-trained monocular depth models, we eliminate the need to rely on per-environment-trained depth networks, enhancing generalization to unseen spaces. We evaluate UnLoc on large-scale synthetic and real-world datasets, demonstrating significant improvements over existing methods in terms of accuracy and robustness. Notably, we achieve $2.7$ times higher localization recall on long sequences (100 frames) and $42.2$ times higher on short ones (15 frames) than the state of the art on the challenging LaMAR HGE dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。