用自然语言描述定位街景位置,提升导航与应急响应效率
Where am I? Cross-View Geo-localization with Natural Language Descriptions
- 通过大模型生成带定位信息的场景文本描述,构建新数据集
- 提出CrossText2Loc方法,召回率提升10%,支持长文本检索
- 不仅给出匹配结果,还解释检索依据,适合需要可解释性的应用
跨视图地理定位通过匹配街景图像与带地理标签的卫星图像或OSM数据库来识别位置。然而,现有研究多聚焦图像到图像检索,较少关注文本引导检索——这对行人导航和应急响应至关重要。本文提出一项新任务:基于自然语言描述进行跨视图地理定位,旨在根据场景文本描述检索对应卫星图像或OSM数据。为此,我们构建了CVG-Text数据集,从多个城市收集跨视图数据,并利用大模型的标注能力生成高质量、含定位细节的场景文本描述。同时提出一种新型文本检索定位方法CrossText2Loc,召回率提升10%,并具备优异的长文本检索能力。在可解释性方面,该方法不仅输出相似度分数,还提供检索理由。更多信息见 https://yejy53.github.io/CVG-Text/
原文摘要 · Abstract (English)
Cross-view geo-localization identifies the locations of street-view images by matching them with geo-tagged satellite images or OSM. However, most existing studies focus on image-to-image retrieval, with fewer addressing text-guided retrieval, a task vital for applications like pedestrian navigation and emergency response. In this work, we introduce a novel task for cross-view geo-localization with natural language descriptions, which aims to retrieve corresponding satellite images or OSM database based on scene text descriptions. To support this task, we construct the CVG-Text dataset by collecting cross-view data from multiple cities and employing a scene text generation approach that leverages the annotation capabilities of Large Multimodal Models to produce high-quality scene text descriptions with localization details. Additionally, we propose a novel text-based retrieval localization method, CrossText2Loc, which improves recall by 10% and demonstrates excellent long-text retrieval capabilities. In terms of explainability, it not only provides similarity scores but also offers retrieval reasons. More information can be found at https://yejy53.github.io/CVG-Text/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。