提出多目标跨视图地理定位新任务与方法
MOGeo: Beyond One-to-One Cross-View Object Geo-localization
- 提出跨视图多目标地理定位新任务,突破单对象假设
- 构建CMLocation基准数据集,包含双版本数据集
- 方法在真实场景下表现更优,适合复杂地理应用
跨视图物体地理定位(CVOGL)旨在将查询图像中的目标物体定位到对应的卫星图像中。现有方法通常假设查询图像仅含单一物体,这与真实世界中复杂的多物体定位需求不符,难以应用于实际场景。为弥合现实设置与现有任务之间的差距,我们提出新任务——跨视图多物体地理定位(CVMOGL)。为推进该任务,我们首先构建基准CMLocation,包含两个数据集:CMLocation-V1和CMLocation-V2。此外,我们提出一种新型跨视图多物体地理定位方法MOGeo,并在多个主流方法上进行对比。在多种应用场景下的大量实验验证了该方法的有效性。结果表明,在更真实的设定下,跨视图物体地理定位仍具挑战性,亟需进一步研究。
原文摘要 · Abstract (English)
Cross-View Object Geo-Localization (CVOGL) aims to locate an object of interest in a query image within a corresponding satellite image. Existing methods typically assume that the query image contains only a single object, which does not align with the complex, multi-object geo-localization requirements in real-world applications, making them unsuitable for practical scenarios. To bridge the gap between the realistic setting and existing task, we propose a new task, called Cross-View Multi-Object Geo-Localization (CVMOGL). To advance the CVMOGL task, we first construct a benchmark, CMLocation, which includes two datasets: CMLocation-V1 and CMLocation-V2. Furthermore, we propose a novel cross-view multi-object geo-localization method, MOGeo, and benchmark it against existing state-of-the-art methods. Extensive experiments are conducted under various application scenarios to validate the effectiveness of our method. The results demonstrate that cross-view object geo-localization in the more realistic setting remains a challenging problem, encouraging further research in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。