用大模型联合建模无人机与卫星图像关系,提升地理定位精度
SkyLink: A Large Vision-Language Model Driven Re-ranking Framework for Cross-View UAV geolocalization
- 利用大视觉语言模型建模无人机与卫星图的跨视图语义关联
- 引入关系感知损失,软标签使近正样本训练更稳定
- 可插拔框架,适配多种模型,显著提升复杂场景定位效果
跨视图无人机地理定位是一项极具挑战性的大规模图像检索任务,旨在通过匹配无人机查询图像与海量地理标记的卫星图像数据库,确定其地理坐标。现有方法通常为各视图分别学习特征表示,并使用简单启发式策略判断特征相似性,忽略了关键的跨视图关系建模。本文提出SkyLink,一种新型即插即用的重排序框架,首次实现对跨视图关系的联合建模,以增强跨视图无人机地理定位能力。SkyLink利用大视觉语言模型(LVLM)建模无人机与卫星视图间的复杂视觉-语义关系,实现高效跨视图匹配。为进一步优化学习过程,我们设计了关系感知损失,通过软标签提供更精细的监督信号,缓解对近正样本的严苛惩罚,从而提升训练稳定性与模型判别力。在多个基础检索架构和基准数据集上的大量实验表明,SkyLink显著提升现有模型的排序效果,在各类复杂场景中持续取得更优性能。
原文摘要 · Abstract (English)
Cross-view UAV geolocalization is fundamentally a challenging large-scale image retrieval task, aiming to determine the geographic coordinates of Unmanned Aerial Vehicle (UAV) queries by matching them against an extensive geo-tagged satellite image database. Most existing methods learn separate feature representations for each view and determine the final prediction using naive heuristics to assess feature similarity, thereby neglecting to model the crucial cross-view relationships. In this paper, we propose SkyLink, a novel plug-and-play ranking framework that pioneers joint relational modeling of inter-view relationships to enhance cross-view UAV geolocalization. SkyLink leverages a Large Vision-Language Model (LVLM) to model the intricate visual-semantic relationships between UAV and satellite views, facilitating effective cross-view matching. To further refine the learning process, we introduce a relational-aware loss. It leverages soft labels to provide a more nuanced supervision signal, mitigating the harsh penalty on near-positive pairs. This approach enhances both training stability and the model's discriminative capacity. Extensive experiments conducted across multiple base retrieval architectures and benchmark datasets demonstrate that SkyLink significantly boosts the ranking effectiveness of existing models, consistently achieving superior performance in various challenging scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。