通过全局-局部一致性与几何等变性提升跨视角地理定位泛化能力
Enhancing Cross-View Geo-Localization Generalization via Global-Local Consistency and Geometric Equivariance
- 采用E(2)可旋转卷积网络提取稳定特征,应对无人机视角变化
- 构建带虚拟超节点的图结构,实现全局语义向局部区域传播
- 在University-1652和SUES-200上刷新跨域定位性能新纪录
跨视角地理定位(CVGL)旨在匹配从迥异视角拍摄的同一地点图像。尽管近期取得进展,现有方法仍面临两大挑战:(1)在不同无人机朝向和视场导致的剧烈外观变化下保持鲁棒性,影响跨域泛化;(2)建立可靠的对应关系,同时捕捉全局场景语义与细粒度局部细节。本文提出EGS框架,以增强跨域泛化能力。具体而言,引入E(2)-可旋转卷积网络编码器,在旋转与视角偏移下提取稳定可靠特征;进一步构建含虚拟超节点的图结构,该节点连接所有局部节点,实现全局语义聚合并重分配至局部区域,强制全局-局部一致性。在University-1652和SUES-200基准上的大量实验表明,EGS持续取得显著性能提升,建立了跨域CVGL新基准。
原文摘要 · Abstract (English)
Cross-view geo-localization (CVGL) aims to match images of the same location captured from drastically different viewpoints. Despite recent progress, existing methods still face two key challenges: (1) achieving robustness under severe appearance variations induced by diverse UAV orientations and fields of view, which hinders cross-domain generalization, and (2) establishing reliable correspondences that capture both global scene-level semantics and fine-grained local details. In this paper, we propose EGS, a novel CVGL framework designed to enhance cross-domain generalization. Specifically, we introduce an E(2)-Steerable CNN encoder to extract stable and reliable features under rotation and viewpoint shifts. Furthermore, we construct a graph with a virtual super-node that connects to all local nodes, enabling global semantics to be aggregated and redistributed to local regions, thereby enforcing global-local consistency. Extensive experiments on the University-1652 and SUES-200 benchmarks demonstrate that EGS consistently achieves substantial performance gains and establishes a new state of the art in cross-domain CVGL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。