arXiv:2603.13917cs.CV2026-03中稿 · the XXV ISPRS Cong…

评估VPR方法在图像对检索中的表现,助力机器人场景注册与建图。

Evaluation of Visual Place Recognition Methods for Image Pair Retrieval in 3D Vision and Robotics

  • 将VPR用于图像对匹配,提升建图与定位精度。
  • 现代全局描述子在复杂场景中表现优异,但效果因场景而异。
  • 适合做视觉里程计、SLAM和三维重建的前处理模块。

视觉位置识别(VPR)是计算机视觉的核心技术,通常用于定位、建图和导航中的图像检索任务。本文将其视为注册流程的前端图像对检索问题,目标是从两组不重叠的图像中找出最优匹配的图像对,以支持后续的场景注册、SLAM和运动恢复结构(SfM)。我们在三个具有挑战性的数据集上对比评估了最先进的VPR方法:NetVLAD类基线、基于分类的全局描述子(CosPlace、EigenPlaces)、特征混合方法(MixVPR),以及基于基础模型的方法(AnyLoc、SALAD、MegaLoc)。结果表明,现代全局描述子在包含感知混淆和序列不完整等复杂场景下,已越来越适合作为开箱即用的图像对检索模块,但其性能表现出明显的领域依赖性,这对构建鲁棒映射与注册系统的选择至关重要。

原文摘要 · Abstract (English)

Visual Place Recognition (VPR) is a core component in computer vision, typically formulated as an image retrieval task for localization, mapping, and navigation. In this work, we instead study VPR as an image pair retrieval front-end for registration pipelines, where the goal is to find top-matching image pairs between two disjoint image sets for downstream tasks such as scene registration, SLAM, and Structure-from-Motion. We comparatively evaluate state-of-the-art VPR families - NetVLAD-style baselines, classification-based global descriptors (CosPlace, EigenPlaces), feature-mixing (MixVPR), and foundation-model-driven methods (AnyLoc, SALAD, MegaLoc) - on three challenging datasets: object-centric outdoor scenes (Tanks and Temples), indoor RGB-D scans (ScanNet-GS), and autonomous-driving sequences (KITTI). We show that modern global descriptor approaches are increasingly suitable as off-the-shelf image pair retrieval modules in challenging scenarios including perceptual aliasing and incomplete sequences, while exhibiting clear, domain-dependent strengths and weaknesses that are critical when choosing VPR components for robust mapping and registration.

视觉定位图像检索机器人建图深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。