用视觉语言模型独立验证定位匹配,提升机器人导航可靠性
Breaking Déjà Vu: Independent Auditing of Visual Place Recognition through Vision-Language Reasoning

- 引入视觉语言模型联合分析查询与候选图像,实现无需依赖环境数据的匹配验证
- 平均召回率提升13.6%,误接受率降至12%,精度保持95%以上
- 适用于安全关键场景,不依赖特定模型或训练数据,可通用部署
视觉定位识别(VPR)是机器人实现精准定位与长期自主导航的关键技术,常用于同时定位与地图构建(SLAM)中的回环检测。然而,实际部署中需设定图像匹配阈值以平衡精确率与召回率,该阈值通常基于标注验证数据调优并固定使用,当环境变化时因缺乏真实标签而不可靠,尤其在安全关键任务中,错误回环会导致轨迹和地图严重偏差。本文提出视觉定位识别审计(Visual Place Recognition Auditing),一种基于视觉语言模型(VLMs)的独立后检索验证框架,通过联合推理查询与候选图像来评估匹配结果。该方法无需依赖架构特有置信度、数据集相关阈值或部署环境先验知识。我们在六个基准数据集上,采用五种先进VPR方法和四种VLMs进行评估。结果显示,基于VLM的审计使平均召回率@1提升13.6%,误接受率降低至12%,精度维持在95%以上,覆盖率达到75%以上。
原文摘要 · Abstract (English)
Visual place recognition (VPR) is a key enabler of accurate localization and long-term autonomous navigation in robotics applications, such as loop closure detection for simultaneous localisation and mapping (SLAM). However, real-world VPR deployment relies on selecting an image matching threshold that balances precision and recall. These thresholds are typically tuned using labeled validation data and fixed during deployment, making them unreliable under environmental changes where ground truth is unavailable. This is particularly problematic in safety-critical robotics, where accepting a false loop closure can corrupt the estimated trajectory and map. In this work, we introduce Visual Place Recognition Auditing, an independent post-retrieval verification framework that leverages Vision-Language Models (VLMs) to assess retrieved matches by reasoning jointly over query and candidate images. Unlike conventional verification methods, our approach performs instance-level verification without requiring architecture-specific confidence measures, dataset-dependent thresholds, or prior knowledge of the deployment environment. We evaluate our method on six benchmark datasets using five state-of-the-art VPR methods and four VLMs. Results show that VLM-based auditing improves recall@1 by 13.6% on average as compared to state-of-the-art methods while reducing false acceptance rates to 12%, maintaining precision above 95% and coverage above 75%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。