arXiv:2505.17911cs.CVcs.AI2025-05被引 5

通过点击定位目标,实现无人机与卫星图像间的精准物体级地理匹配。

Object-level Cross-view Geo-localization with Location Enhancement and Multi-Head Cross Attention

  • 利用高斯核传递点击位置信息,保持特征网络中定位精度。
  • 在CVOGL数据集上达到当前最佳性能,支持少样本快速泛化。
  • 适合搜救、巡检等需要精确物体定位的实用场景。

跨视角地理定位旨在通过将无人机或地面相机拍摄的查询图像与地理标注的卫星图像匹配,确定其位置。传统方法仅关注图像级别定位,但搜救、基础设施巡检和精准配送等应用需物体级别精度,使用户可点击无人机图像中的特定物体,获取其精确地理标签信息。然而,视角、时间及成像条件差异,尤其在大量卫星图像中识别视觉相似物体时,带来巨大挑战。为此,我们提出物体级跨视角地理定位网络(OCGNet),通过高斯核转移(GKT)整合用户指定的点击位置,确保整个网络中定位信息的保留。该提示同时嵌入特征编码器和特征匹配模块,保障对象特异性定位鲁棒性。此外,OCGNet引入位置增强(LE)模块与多头交叉注意力(MHCA)模块,可自适应强调对象特征或扩展关注至相关上下文区域。OCGNet在公开数据集CVOGL上达到当前最优表现,并展现出少样本学习能力,能从有限样本有效泛化,适用于多样化应用场景(https://github.com/ZheyangH/OCGNet)。

原文摘要 · Abstract (English)

Cross-view geo-localization determines the location of a query image, captured by a drone or ground-based camera, by matching it to a geo-referenced satellite image. While traditional approaches focus on image-level localization, many applications, such as search-and-rescue, infrastructure inspection, and precision delivery, demand object-level accuracy. This enables users to prompt a specific object with a single click on a drone image to retrieve precise geo-tagged information of the object. However, variations in viewpoints, timing, and imaging conditions pose significant challenges, especially when identifying visually similar objects in extensive satellite imagery. To address these challenges, we propose an Object-level Cross-view Geo-localization Network (OCGNet). It integrates user-specified click locations using Gaussian Kernel Transfer (GKT) to preserve location information throughout the network. This cue is dually embedded into the feature encoder and feature matching blocks, ensuring robust object-specific localization. Additionally, OCGNet incorporates a Location Enhancement (LE) module and a Multi-Head Cross Attention (MHCA) module to adaptively emphasize object-specific features or expand focus to relevant contextual regions when necessary. OCGNet achieves state-of-the-art performance on a public dataset, CVOGL. It also demonstrates few-shot learning capabilities, effectively generalizing from limited examples, making it suitable for diverse applications (https://github.com/ZheyangH/OCGNet).

地理定位跨视角匹配少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。