用多视角图像集提升地理定位精度,更贴近人类观察方式。
Cross-View Image Set Geo-Localization
- 以多张不同视角的图片作为查询集,增强定位信息多样性。
- 在SetVL-480K数据集上定位准确率提升超22%。
- 适用于单图、序列图和多视角集,灵活性强,适合导航与增强现实应用。
跨视角地理定位(CVGL)广泛应用于机器人导航和增强现实等领域。现有方法通常使用单张图像或固定视角图像序列作为查询,限制了视角多样性。而人类在视觉判断位置时,常通过移动获取多个视角。这提示融合多样视觉线索可提升定位可靠性。为此,我们提出新任务:跨视角图像集地理定位(Set-CVGL),即以多张不同视角的图像作为查询集进行定位。为支持该任务,我们构建了SetVL-480K基准数据集,包含48万张全球范围采集的地面图像及其对应卫星图像,每张卫星图像平均关联40张不同视角和位置的地面图像。此外,我们提出FlexGeo方法,专为Set-CVGL设计,也可适配单图和图像序列输入。其核心模块包括:相似性引导特征融合器(SFF),无需先验内容依赖即可自适应融合特征;个体级属性学习器(IAL),利用每张图像的地理属性实现全面场景感知。FlexGeo在SetVL-480K及两个公开数据集SeqGeo和KITTI-CVL上均优于现有方法,在SetVL-480K上定位准确率提升超过22%。
原文摘要 · Abstract (English)
Cross-view geo-localization (CVGL) has been widely applied in fields such as robotic navigation and augmented reality. Existing approaches primarily use single images or fixed-view image sequences as queries, which limits perspective diversity. In contrast, when humans determine their location visually, they typically move around to gather multiple perspectives. This behavior suggests that integrating diverse visual cues can improve geo-localization reliability. Therefore, we propose a novel task: Cross-View Image Set Geo-Localization (Set-CVGL), which gathers multiple images with diverse perspectives as a query set for localization. To support this task, we introduce SetVL-480K, a benchmark comprising 480,000 ground images captured worldwide and their corresponding satellite images, with each satellite image corresponds to an average of 40 ground images from varied perspectives and locations. Furthermore, we propose FlexGeo, a flexible method designed for Set-CVGL that can also adapt to single-image and image-sequence inputs. FlexGeo includes two key modules: the Similarity-guided Feature Fuser (SFF), which adaptively fuses image features without prior content dependency, and the Individual-level Attributes Learner (IAL), leveraging geo-attributes of each image for comprehensive scene perception. FlexGeo consistently outperforms existing methods on SetVL-480K and two public datasets, SeqGeo and KITTI-CVL, achieving a localization accuracy improvement of over 22% on SetVL-480K.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。