arXiv:2603.25819cs.CV2026-03被引 4

利用3D几何先验实现地面与航拍图像的精准定位与双向生成

Geo$^\textbf{2}$: Geometry-Guided Cross-view Geo-Localization and Image Synthesis

  • 构建共享3D感知潜空间,缓解地面与航拍视角差异
  • 在CVUSA/CVACT/VIGOR上定位精度达92.1%/85.6%/87.3%,合成效果更自然
  • 适合从事地理视觉、跨视角生成的研究者使用

跨视角地理空间学习包含两个核心任务:跨视角地理定位(CVGL)和跨视角图像合成(CVIS),二者均依赖于地面与航拍图像间的几何对应关系。现有几何基础模型(如VGGT)虽具备提取通用3D几何特征的能力,但在跨视角地理空间任务中的潜力尚未充分挖掘。本文提出统一框架Geo²,利用几何先验(GFMs)联合完成地理定位与双向图像合成。针对地面与航拍视角差异大导致直接应用困难的问题,提出GeoMap将地面与航拍特征嵌入共享3D感知潜空间,有效减少跨视角偏差,并自然支持双向图像合成。进一步设计基于几何感知嵌入的流匹配模型GeoFlow,引入一致性损失约束双向合成的潜空间对齐,确保双向一致性。在标准基准(CVUSA、CVACT、VIGOR)上的大量实验表明,Geo²在定位与合成任务上均达到最先进性能,验证了3D几何先验在跨视角地理空间学习中的有效性。

原文摘要 · Abstract (English)

Cross-view geo-spatial learning consists of two important tasks: Cross-View Geo-Localization (CVGL) and Cross-View Image Synthesis (CVIS), both of which rely on establishing geometric correspondences between ground and aerial views. Recent Geometric Foundation Models (GFMs) have demonstrated strong capabilities in extracting generalizable 3D geometric features from images, but their potential in cross-view geo-spatial tasks remains underexplored. In this work, we present Geo^2, a unified framework that leverages Geometric priors from GFMs (e.g., VGGT) to jointly perform geo-spatial tasks, CVGL and bidirectional CVIS. Despite the 3D reconstruction ability of GFMs, directly applying them to CVGL and CVIS remains challenging due to the large viewpoint gap between ground and aerial imagery. We propose GeoMap, which embeds ground and aerial features into a shared 3D-aware latent space, effectively reducing cross-view discrepancies for localization. This shared latent space naturally bridges cross-view image synthesis in both directions. To exploit this, we propose GeoFlow, a flow-matching model conditioned on geometry-aware latent embeddings. We further introduce a consistency loss to enforce latent alignment between the two synthesis directions, ensuring bidirectional coherence. Extensive experiments on standard benchmarks, including CVUSA, CVACT, and VIGOR, demonstrate that Geo^2 achieves state-of-the-art performance in both localization and synthesis, highlighting the effectiveness of 3D geometric priors for cross-view geo-spatial learning.

地理定位图像合成3D几何跨视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。