arXiv:2608.09321cs.CV2026-08

不依赖几何扭曲,直接在特征空间挖掘跨视角语义共识,提升街景与卫星图像定位精度。

Warp-free Cross-view Geo-localization via Feature-space Consensus Mining

论文配图:Warp-free Cross-view Geo-localization via Feature-space Consensus Mining
图 1 · 摘自论文原文
  • 在特征空间动态挖掘跨视角语义共识,避免显式几何扭曲带来的失真。
  • 在四个基准上达到最先进性能,显著提升跨视角定位准确率。
  • 适合需要高鲁棒性地理定位的视觉检索、自动驾驶等应用。

跨视角地理定位因街景与卫星影像间视角剧变和外观差异大而极具挑战。现有方法常依赖几何扭曲以暴露共可见线索,但此类变换受限于空间假设,在视点依赖可见性下会引入严重视觉畸变,导致噪声监督和脆弱对应关系。为此,我们提出一种无需显式几何扭曲的联合视角共识引导学习框架。通过训练时的辅助联合视角路径,实现跨视角直接交互,使每视角选择性聚合支持性证据,形成统一共识表示。为解决单视角与联合视角流间的特征异质性,引入全局模式探针作为语义词典,将不同模态投影至严格对齐的度量空间。在共识引导的对比目标指导下,单视角嵌入被显式拉向联合视角锚点,将共识挖掘能力蒸馏至单视角编码器,从而在推理时实现鲁棒检索。大量实验表明,该方法在四个标准基准上均取得最先进性能,凸显发现跨视角语义共识对可靠地理定位的重要性。

原文摘要 · Abstract (English)

Cross-view geo-localization is challenging due to drastic viewpoint changes and large appearance discrepancies between street-level and satellite imagery. Although existing methods often use geometric warping to expose co-visible cues, such transformations rely on restrictive spatial assumptions and inevitably introduce severe visual distortions under view-dependent visibility, yielding noisy supervision and fragile correspondences. To overcome this, we propose a novel joint-view consensus-guided learning framework that entirely bypasses explicit geometric warping. Instead of forcing rigid spatial alignment, we dynamically mine and adaptively strengthen a semantic consensus directly within the feature space. Specifically, an auxiliary joint-view pathway during training enables direct cross-view interaction, allowing each view to selectively aggregate corroborative evidence into a unified consensus representation. To resolve feature heterogeneity among the single- and joint-view streams, we introduce global pattern probes acting as a semantic dictionary to project divergent modalities into a strictly aligned metric space. Guided by a consensus-mediated contrastive objective, single-view embeddings are explicitly pulled toward the joint-view anchor during training, distilling this consensus-mining capability into the single-view encoders for robust retrieval at inference. Extensive experiments demonstrate that our method achieves state-of-the-art performance across four standard benchmarks, underscoring the importance of discovering cross-view semantic consensus for reliable geo-localization.

地理定位特征对齐跨视角检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。