arXiv:2503.06840cs.CV2025-03中稿 · the IEEE/RSJ Inter…被引 4

通过预测序列匹配可信度,提升视觉定位在复杂场景下的准确性

Improving Visual Place Recognition with Sequence-Matching Receptiveness Prediction

  • 训练模型预测每帧序列匹配的可信度,决定是否信任序列结果
  • 在3个基准数据集上显著提升7种主流VPR方法的性能
  • 适用于各类传统与先进视觉定位方法,增强系统鲁棒性

在视觉定位(VPR)中,融合图像序列的时间信息可提升复杂场景下的表现。然而,现有过滤与序列匹配方法的效果难以预测,有时反而降低性能。本文提出一种新的监督学习方法,用于预测每帧序列匹配的可信度(SMR),使系统能智能选择是否采纳序列匹配输出。该方法不依赖具体VPR技术,可有效预测SMR,显著提升多种经典与先进VPR方法(包括CosPlace、MixVPR、EigenPlaces、SALAD、AP-GeM、NetVLAD和SAD)在三个基准数据集(Nordland、Oxford RobotCar、SFU-Mountain)上的表现。我们还探索了用预测器替代被丢弃匹配的互补策略,并通过消融实验分析了预测器与序列长度之间的交互关系。

原文摘要 · Abstract (English)

In visual place recognition (VPR), filtering and sequence-based matching approaches can improve performance by integrating temporal information across image sequences, especially in challenging conditions. While these methods are commonly applied, their effects on system behavior can be unpredictable and can actually make performance worse in certain situations. In this work, we present a new supervised learning approach that learns to predict the per-frame sequence matching receptiveness (SMR) of VPR techniques, enabling the system to selectively decide when to trust the output of a sequence matching system. Our approach is agnostic to the underlying VPR technique and effectively predicts SMR, and hence significantly improves VPR performance across a large range of state-of-the-art and classical VPR techniques (namely CosPlace, MixVPR, EigenPlaces, SALAD, AP-GeM, NetVLAD and SAD), and across three benchmark VPR datasets (Nordland, Oxford RobotCar, and SFU-Mountain). We also provide insights into a complementary approach that uses the predictor to replace discarded matches, and present ablation studies including an analysis of the interactions between our SMR predictor and the selected sequence length.

视觉定位序列匹配可信度预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。