arXiv:2412.06153cs.CV2024-12ICCV被引 2

用高维签名融合多条件图像,实现高效精准的视觉定位

A Hyperdimensional One Place Signature to Represent Them All: Stackable Descriptors For Visual Place Recognition

  • 将不同环境下的参考图像特征融合为高维签名
  • 在多个数据集上显著提升召回率,且计算开销几乎不变
  • 适合需要大规模、多场景定位的应用开发者

视觉位置识别(VPR)通过比对查询图像与带地理标签的参考图像数据库,实现粗略定位。近年来深度学习架构和训练方法的进步提升了模型对环境变化的鲁棒性,但训练和匹配计算量随环境种类增加而增长。本文提出高维单地签名(HOPS),通过融合在不同条件下采集的多个参考集特征,同时提升性能、降低计算成本并增强可扩展性。HOPS利用高维计算框架,可无缝扩展至任意数量的环境条件。大量实验表明,该方法在所有评估的VPR方法和数据集上均显著提升召回率。无需额外计算代价即可任意融合参考图像,拓展出多种新应用:特征维度压缩无性能损失、堆叠合成图像、对整个路线或环境段进行粗略定位。

原文摘要 · Abstract (English)

Visual Place Recognition (VPR) enables coarse localization by comparing query images to a reference database of geo-tagged images. Recent breakthroughs in deep learning architectures and training regimes have led to methods with improved robustness to factors like environment appearance change, but with the downside that the required training and/or matching compute scales with the number of distinct environmental conditions encountered. Here, we propose Hyperdimensional One Place Signatures (HOPS) to simultaneously improve the performance, compute and scalability of these state-of-the-art approaches by fusing the descriptors from multiple reference sets captured under different conditions. HOPS scales to any number of environmental conditions by leveraging the Hyperdimensional Computing framework. Extensive evaluations demonstrate that our approach is highly generalizable and consistently improves recall performance across all evaluated VPR methods and datasets by large margins. Arbitrarily fusing reference images without compute penalty enables numerous other useful possibilities, three of which we demonstrate here: descriptor dimensionality reduction with no performance penalty, stacking synthetic images, and coarse localization to an entire traverse or environmental section.

视觉定位高维计算特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。