用体素渲染生成新视角特征,提升相机重定位鲁棒性
FaVoR: Features via Voxel Rendering for Camera Relocalization
- 构建稀疏体素地图,通过体素渲染合成图像特征描述子
- 在7-Scenes数据集上翻译误差中位数降低39%、显著优于现有方法
- 适合对实时性与内存敏感的室内场景重定位应用
相机重定位方法从密集图像配准到直接从查询图像回归相机位姿不等。其中,稀疏特征匹配因其高效、通用且轻量而广受欢迎。然而,基于特征的方法在视角和外观变化显著时易出现匹配失败,导致位姿估计不准。为此,我们提出一种新方法:利用全局稀疏但局部稠密的2D特征三维表示。通过在多帧序列中追踪并三角化关键点,构建一个优化用于渲染跟踪过程中观察到的图像块描述子的稀疏体素地图。给定初始位姿估计后,首先使用体素渲染合成描述子,再进行特征匹配以估计相机位姿。该方法可生成未见视角的描述子,增强对视图变化的鲁棒性。我们在7-Scenes和Cambridge Landmarks数据集上进行了广泛评估。结果表明,在室内环境中,本方法显著优于现有最先进特征表示技术,翻译误差中位数最高提升39%;在室外场景中性能与其它方法相当,同时保持更低的内存与计算开销。
原文摘要 · Abstract (English)
Camera relocalization methods range from dense image alignment to direct camera pose regression from a query image. Among these, sparse feature matching stands out as an efficient, versatile, and generally lightweight approach with numerous applications. However, feature-based methods often struggle with significant viewpoint and appearance changes, leading to matching failures and inaccurate pose estimates. To overcome this limitation, we propose a novel approach that leverages a globally sparse yet locally dense 3D representation of 2D features. By tracking and triangulating landmarks over a sequence of frames, we construct a sparse voxel map optimized to render image patch descriptors observed during tracking. Given an initial pose estimate, we first synthesize descriptors from the voxels using volumetric rendering and then perform feature matching to estimate the camera pose. This methodology enables the generation of descriptors for unseen views, enhancing robustness to view changes. We extensively evaluate our method on the 7-Scenes and Cambridge Landmarks datasets. Our results show that our method significantly outperforms existing state-of-the-art feature representation techniques in indoor environments, achieving up to a 39% improvement in median translation error. Additionally, our approach yields comparable results to other methods for outdoor scenarios while maintaining lower memory and computational costs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。