提出新相似度度量,提升视觉语言联合定位精度
MAPS: Multi-Anchor Projection Similarity for Joint Vision-Language Geo-Localization

- 用多锚点投影相似度构建联合语义空间
- 在Geo-Localization任务上达到最新最佳性能
- 适合需要跨模态定位的智能导航系统
人类通过融合视觉感知与语言语义来定位地点,形成直观而结构化的场景理解。现有地理定位模型虽在跨视图、跨模态设置中取得进展,但大多依赖点对点对齐,难以应对视觉-语言联合查询。此类查询中,视觉与文本线索并非独立参考,而是共同定义目标定位的语义子空间。本文将视觉-语言地理定位(VLGL)建模为多锚点几何对齐问题,提出统一框架。核心是多锚点投影相似度(MAPS),在高维空间中由视觉与文本查询特征构建锚平面,并以目标特征在该平面上的投影长度衡量相似性。相比仅评估成对关系的余弦相似度,MAPS捕捉目标特征与联合查询子空间间的几何一致性,提供更优检索排序。为此设计基于MAPS的对比损失,引导目标特征向对应锚平面靠拢。该框架、度量与训练目标协同实现VLGL领域最先进性能。
原文摘要 · Abstract (English)
Humans localize places by integrating perceptual cues from vision with semantic reasoning from language, forming a scene understanding that is both intuitive and structured. Although existing geo-localization models have made substantial progress in cross-view and cross-modal settings, they are largely built upon point-to-point alignment, which is insufficient for joint vision-language queries. In such queries, visual and textual cues do not simply act as independent references, but jointly define a semantic subspace for locating the target. In this paper, we formulate vision-language geo-localization (VLGL) with joint image-text queries as a multi-anchor geometric alignment problem and propose a unified framework for this setting. To realize this formulation, we propose Multi-Anchor Projection Similarity (MAPS), a new metric which constructs an anchor plane from visual and textual query features in a high-dimensional space and measures similarity by the projection length of the target feature onto this plane. Unlike cosine similarity which evaluates isolated pairwise relations, MAPS captures the geometric consistency between the target feature and the joint query subspace, providing a more discriminative ranking criterion during retrieval. To make the learned representation consistent with this geometry, we further introduce a MAPS-based contrastive loss that drives target features toward the corresponding anchor plane. The proposed framework, similarity metric, and training objective jointly yield state-of-the-art performance in VLGL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。