arXiv:2504.08901cs.CV2025-04被引 1

用NeRF提升单目相机定位精度,误差低于0.03米。

HAL-NeRF: High Accuracy Localization Leveraging Neural Radiance Fields

  • 结合CNN回归与粒子滤波优化,实现高精度定位
  • 在7-Scenes和Cambridge数据集上误差分别低至0.025米和0.04米
  • 适合需要高精度定位的XR与机器人应用

精确的相机定位在扩展现实(XR)和机器人领域至关重要。仅依赖相机捕获进行定位是一种低成本方案,可在室内外大场景中实现定位,但难以达到高精度。现有方法如绝对位姿回归(APR)在室外场景中的中位平移误差超过0.5米。本文提出HAL-NeRF,将CNN位姿回归器与基于蒙特卡洛粒子滤波的精修模块相结合。采用Nerfacto模型(一种NeRF实现)生成高质量新视角图像,用于训练回归器并计算粒子滤波中的光度损失。该方法显著提升了定位性能。在7-Scenes和Cambridge Landmarks数据集上,平均中位误差分别为0.025米、0.59度和0.04米、0.58度,代价是计算时间增加。结果表明,将APR与基于NeRF的精修结合可大幅提升单目重定位精度。

原文摘要 · Abstract (English)

Precise camera localization is a critical task in XR applications and robotics. Using only the camera captures as input to a system is an inexpensive option that enables localization in large indoor and outdoor environments, but it presents challenges in achieving high accuracy. Specifically, camera relocalization methods, such as Absolute Pose Regression (APR), can localize cameras with a median translation error of more than $0.5m$ in outdoor scenes. This paper presents HAL-NeRF, a high-accuracy localization method that combines a CNN pose regressor with a refinement module based on a Monte Carlo particle filter. The Nerfacto model, an implementation of Neural Radiance Fields (NeRFs), is used to augment the data for training the pose regressor and to measure photometric loss in the particle filter refinement module. HAL-NeRF leverages Nerfacto's ability to synthesize high-quality novel views, significantly improving the performance of the localization pipeline. HAL-NeRF achieves state-of-the-art results that are conventionally measured as the average of the median per scene errors. The translation error was $0.025m$ and the rotation error was $0.59$ degrees and 0.04m and 0.58 degrees on the 7-Scenes dataset and Cambridge Landmarks datasets respectively, with the trade-off of increased computational time. This work highlights the potential of combining APR with NeRF-based refinement techniques to advance monocular camera relocalization accuracy.

相机定位NeRF姿态估计深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。