通过区分环境变化与地点身份,提升多参考图像的定位精度
DisPlace: Discriminative Place Projections for Multi-Reference Visual Place Recognition

- 将多参考描述符融合建模为广义特征值问题,增强不同地点间的可区分性
- 在6个数据集上49/54种条件下超越7种基线方法,尤其在视角和无结构场景表现更优
- 仅需更少存储空间,适合部署在资源受限的实时定位系统中
视觉位置识别(VPR)的核心挑战在于如何在不同环境条件和视角下匹配查询图像与参考地图。尽管多参考遍历能提升鲁棒性,现有融合策略或均匀聚合参考信息,或依赖启发式选择,无法区分保持稳定地点身份的描述符变化与由环境或视角变化引起的差异。本文提出DisPlace,一种多参考VPR框架,将多个参考描述符融合为单一紧凑且具有判别力的位置表示。DisPlace将描述符融合建模为广义特征值问题,最大化不同地点间的分离度,同时抑制同一地点在不同参考间的内部变异,而非保留整体描述符方差。相比现有方法,DisPlace利用多参考遍历中的变化,识别出哪些描述符维度组合能保留地点身份,哪些捕捉条件或视角特异性变化。我们在Oxford RobotCar、Nordland、Pittsburgh30k和Google Landmarks v2数据集上,使用六种先进的VPR描述符进行评估。DisPlace在54种外观变化条件下有49次优于7种多参考基线,在视角和非结构化设置下持续提升描述符级融合性能,且推理时所需存储空间小于所有对比方法。
原文摘要 · Abstract (English)
A key challenge in Visual Place Recognition (VPR) is matching query images against reference maps captured under diverse environmental conditions and viewpoints. While multiple reference traversals improve robustness, existing fusion strategies either aggregate references uniformly or rely on heuristic selection, without distinguishing descriptor variations that preserve stable place identity from those caused by changing conditions or viewpoints. In this paper, we propose DisPlace, a multi-reference VPR framework that fuses multiple reference descriptors into a single compact and discriminative place representation. DisPlace formulates descriptor fusion as a generalized eigenvalue problem that maximizes between-place separability while suppressing within-place variation across references, rather than preserving overall descriptor variance. Unlike existing multi-reference fusion methods, DisPlace exploits variation across reference traversals to identify which linear combinations of descriptor dimensions preserve place identity and which capture condition- or viewpoint-specific variation. We evaluate DisPlace on Oxford RobotCar, Nordland, Pittsburgh30k, and Google Landmarks v2 across six state-of-the-art VPR descriptors. DisPlace outperforms seven multi-reference baselines in 49 out of 54 appearance-varying conditions, consistently improves descriptor-level fusion performance under viewpoint and unstructured settings, and requires less storage during inference than all compared fusion methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。