测试时用地图数据微调模型,显著提升视觉定位准确率
The Overlooked Value of Test-time Reference Sets in Visual Place Recognition
- 测试时利用目标域地图数据微调模型,缓解训练测试分布差异
- 在挑战性数据集上平均召回率提升2.3%(Recall@1)
- 适用于真实场景中的视觉定位任务,尤其适合部署时有地图信息的系统
给定查询图像,视觉定位(VPR)的目标是从参考数据库中检索出同一地点的图像,且对视角和外观变化具有鲁棒性。近期研究表明,一些VPR基准已可通过使用视觉基础模型主干并基于大规模、多样化的特定VPR数据集训练的方法解决。然而,在测试环境与常规训练数据差异较大的情况下,仍存在若干基准极具挑战性。本文提出一种未被充分重视的信息源——测试时的参考集(即“地图”),该地图包含目标域的图像与位姿信息,在部分VPR应用中可提前获取。我们提出在测试时对模型进行简单的参考集微调(RSF),显著提升了SOTA方法在这些挑战性数据集上的表现(平均Recall@1提升约2.3%)。微调后模型仍保持良好泛化能力,且在多种不同测试数据集上均有效。
原文摘要 · Abstract (English)
Given a query image, Visual Place Recognition (VPR) is the task of retrieving an image of the same place from a reference database with robustness to viewpoint and appearance changes. Recent works show that some VPR benchmarks are solved by methods using Vision-Foundation-Model backbones and trained on large-scale and diverse VPR-specific datasets. Several benchmarks remain challenging, particularly when the test environments differ significantly from the usual VPR training datasets. We propose a complementary, unexplored source of information to bridge the train-test domain gap, which can further improve the performance of State-of-the-Art (SOTA) VPR methods on such challenging benchmarks. Concretely, we identify that the test-time reference set, the "map", contains images and poses of the target domain, and must be available before the test-time query is received in several VPR applications. Therefore, we propose to perform simple Reference-Set-Finetuning (RSF) of VPR models on the map, boosting the SOTA (~2.3% increase on average for Recall@1) on these challenging datasets. Finetuned models retain generalization, and RSF works across diverse test datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。