arXiv:2409.15514cs.CV2024-09被引 5

用图结构建模城市地理定位,提升跨视角检索精度。

SpaGBOL: Spatial-Graph-Based Orientated Localisation

  • 构建首个面向跨视角定位的图结构数据集,每节点含多张街景图。
  • 引入图神经网络,利用邻近节点特征相似性提升定位准确率。
  • 提出基于邻域方位向量的过滤方法,适合城市导航与自动驾驶场景。

城市区域内的跨视角地理定位面临数据集与技术缺乏地理空间结构的问题。本文提出使用图表示来建模局部观测序列及其目标位置的连接关系。将数据建模为图后,可通过新参数配置采样生成未见序列。为利用这一新增信息,设计基于图神经网络的架构,生成空间强嵌入,增强孤立图像嵌入的区分能力。提出SpaGBOL,包含三项创新:1)首个用于跨视角地理定位的图结构数据集,每个节点包含多张街景图像,提升泛化能力;2)首次将图神经网络引入该任务,构建了利用节点邻近性与特征相似性相关性的系统;3)利用图表示的独特性质,提出一种基于邻域方位向量的新型检索过滤方法。在未见测试图上,SpaGBOL达到当前最优性能,相较以往方法相对提升11%的Top-1检索准确率,结合方位向量匹配过滤后提升达50%。

原文摘要 · Abstract (English)

Cross-View Geo-Localisation within urban regions is challenging in part due to the lack of geo-spatial structuring within current datasets and techniques. We propose utilising graph representations to model sequences of local observations and the connectivity of the target location. Modelling as a graph enables generating previously unseen sequences by sampling with new parameter configurations. To leverage this newly available information, we propose a GNN-based architecture, producing spatially strong embeddings and improving discriminability over isolated image embeddings. We outline SpaGBOL, introducing three novel contributions. 1) The first graph-structured dataset for Cross-View Geo-Localisation, containing multiple streetview images per node to improve generalisation. 2) Introducing GNNs to the problem, we develop the first system that exploits the correlation between node proximity and feature similarity. 3) Leveraging the unique properties of the graph representation - we demonstrate a novel retrieval filtering approach based on neighbourhood bearings. SpaGBOL achieves state-of-the-art accuracies on the unseen test graph - with relative Top-1 retrieval improvements on previous techniques of 11%, and 50% when filtering with Bearing Vector Matching on the SpaGBOL dataset.

地理定位图神经网络街景检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。