arXiv:2504.11134cs.CV2025-04中稿 · Scandinavian Confe…被引 1

用非视觉信息重排图像检索结果,提升定位精度。

Visual Re-Ranking with Non-Visual Side Information

  • 基于图神经网络,融合视觉与非视觉数据重排检索结果
  • 在室外和室内场景中,定位准确率显著提升
  • 适合已有传感器数据的智能导航系统使用

视觉地点识别通常采用全局图像描述符检索最相似的数据库图像,再通过重排序方法优化结果。然而现有方法仅依赖初始检索所用的视觉描述符,信号有限。本文提出广义上下文相似性聚合(GCSA),一种基于图神经网络的重排序方法,可同时利用视觉描述符及其他可用的辅助信息,如周边WiFi或蓝牙信号强度、数据库图像的相机位姿等。这些信息在许多实际应用中已存在或获取成本低。模型通过亲和向量实现异构多模态输入的共享编码。在覆盖室内外场景的两个大规模数据集上进行训练与评估,实验表明不仅在图像检索指标上表现更优,对下游视觉定位任务也有显著提升。

原文摘要 · Abstract (English)

The standard approach for visual place recognition is to use global image descriptors to retrieve the most similar database images for a given query image. The results can then be further improved with re-ranking methods that re-order the top scoring images. However, existing methods focus on re-ranking based on the same image descriptors that were used for the initial retrieval, which we argue provides limited additional signal. In this work we propose Generalized Contextual Similarity Aggregation (GCSA), which is a graph neural network-based re-ranking method that, in addition to the visual descriptors, can leverage other types of available side information. This can for example be other sensor data (such as signal strength of nearby WiFi or BlueTooth endpoints) or geometric properties such as camera poses for database images. In many applications this information is already present or can be acquired with low effort. Our architecture leverages the concept of affinity vectors to allow for a shared encoding of the heterogeneous multi-modal input. Two large-scale datasets, covering both outdoor and indoor localization scenarios, are utilized for training and evaluation. In experiments we show significant improvement not only on image retrieval metrics, but also for the downstream visual localization task.

视觉定位重排序多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。