arXiv:2503.08170cs.CV2025-03

通过上下文查询增强地标识别,提升复杂场景下的图像定位精度

CQVPR: Landmark-aware Contextual Queries for Visual Place Recognition

  • 用可学习的上下文查询自动捕捉地标及其周边环境特征
  • 在多个数据集上优于现有方法,尤其在城市复杂场景中表现突出
  • 适合需要高精度定位的自动驾驶与地图构建应用

视觉位置识别(VPR)旨在根据给定查询图像,在带有地理标签的图像数据库中估计其位置。准确识别图像中的位置依赖于地标检测。然而,在城市环境中存在大量视觉相似的现代建筑等地标,因此不仅需利用地标信息,还需考虑其周围上下文,如树木、道路等特征。本文提出上下文查询视觉位置识别(CQVPR),将上下文信息与像素级视觉特征相结合。通过一组可学习的上下文查询,模型能自动学习与地标及其周围区域相关的高层语义上下文。每个查询对应的注意力热图作为上下文感知特征,为场景理解提供有效线索。此外,设计了查询匹配损失以监督上下文查询的提取过程。在多个数据集上的大量实验表明,该方法在挑战性场景下显著优于现有先进方法。

原文摘要 · Abstract (English)

Visual Place Recognition (VPR) aims to estimate the location of the given query image within a database of geo-tagged images. To identify the exact location in an image, detecting landmarks is crucial. However, in some scenarios, such as urban environments, there are numerous landmarks, such as various modern buildings, and the landmarks in different cities often exhibit high visual similarity. Therefore, it is essential not only to leverage the landmarks but also to consider the contextual information surrounding them, such as whether there are trees, roads, or other features around the landmarks. We propose the Contextual Query VPR (CQVPR), which integrates contextual information with detailed pixel-level visual features. By leveraging a set of learnable contextual queries, our method automatically learns the high-level contexts with respect to landmarks and their surrounding areas. Heatmaps depicting regions that each query attends to serve as context-aware features, offering cues that could enhance the understanding of each scene. We further propose a query matching loss to supervise the extraction process of contextual queries. Extensive experiments on several datasets demonstrate that the proposed method outperforms other state-of-the-art methods, especially in challenging scenarios.

视觉定位上下文建模地标识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。