arXiv:2511.14109cs.CV2025-11被引 3

提出不对称聚合方法提升视觉位置识别准确率

A2GC: Asymmetric Aggregation with Geometric Constraints for Locally Aggregated Descriptors

  • 采用行列归一化与独立边际校准实现特征与聚类中心的非对称匹配
  • 在MSLS、NordLand、Pittsburgh数据集上均优于现有方法
  • 通过可学习坐标嵌入增强空间感知,适合高精度定位场景

视觉位置识别(VPR)旨在利用视觉线索将查询图像与数据库进行匹配。当前先进方法通过聚合深度主干网络的特征生成全局描述符。基于最优传输的聚合方法将特征到聚类的分配重构为运输问题,但标准Sinkhorn算法对源和目标边际采取对称处理,在图像特征与聚类中心分布差异显著时效果受限。本文提出一种带几何约束的非对称聚合方法A²GC-VPR,通过行-列归一化平均与独立边际校准,实现适应分布差异的非对称匹配。引入可学习坐标嵌入,计算融合特征相似性的兼容性分数,促使空间邻近特征被分配至同一聚类,增强空间感知能力。在MSLS、NordLand和Pittsburgh数据集上的实验结果表明,该方法显著提升了匹配准确率与鲁棒性。

原文摘要 · Abstract (English)

Visual Place Recognition (VPR) aims to match query images against a database using visual cues. State-of-the-art methods aggregate features from deep backbones to form global descriptors. Optimal transport-based aggregation methods reformulate feature-to-cluster assignment as a transport problem, but the standard Sinkhorn algorithm symmetrically treats source and target marginals, limiting effectiveness when image features and cluster centers exhibit substantially different distributions. We propose an asymmetric aggregation VPR method with geometric constraints for locally aggregated descriptors, called $A^2$GC-VPR. Our method employs row-column normalization averaging with separate marginal calibration, enabling asymmetric matching that adapts to distributional discrepancies in visual place recognition. Geometric constraints are incorporated through learnable coordinate embeddings, computing compatibility scores fused with feature similarities, thereby promoting spatially proximal features to the same cluster and enhancing spatial awareness. Experimental results on MSLS, NordLand, and Pittsburgh datasets demonstrate superior performance, validating the effectiveness of our approach in improving matching accuracy and robustness.

视觉定位特征聚合最优传输空间感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。