用大模型让文字描述更准定位无人机拍摄位置。
UniGeo: A Multi-modal Large Language Model for Text-Guided Cross-View Geo-Localization

- 统一框架融合语义理解、视角转换和候选验证
- 在GeoText-1652上提升R@10 13.59点,mAP 2.83点
- 适合需要精确地理定位的遥感与导航应用
文本引导的无人机地理定位旨在从大规模图像库中根据自然语言描述识别目标区域。现有方法多将此任务视为开放文本与候选图像间的直接匹配,但不完整查询和高度相似候选常导致全局跨模态匹配不足,难以实现精细定位。本文提出UniGeo,一种统一的多模态大语言模型,用于文本引导的无人机地理定位。基于共享视觉-语言框架,UniGeo同时支持地理语义理解、跨视角语义生成和候选级别验证。具体而言,通过地理语义学习建立局部场景元素、空间关系与语言描述之间的稳定对应,并进一步通过跨视角生成建模无人机与卫星视图间的语义映射。基于这些能力,一个即插即用的验证模块对高度混淆的候选进行细粒度区分。我们还引入多阶段训练策略,逐步学习地理语义理解、跨视角生成与候选验证,提升对文本引导地理定位任务的适应性。实验表明,在多个检索骨干网络上均取得一致提升。在GeoText-1652数据集上,UniGeo使R@10和mAP分别提升13.59和2.83个百分点,验证了其在细粒度文本引导无人机地理定位中的有效性。
原文摘要 · Abstract (English)
Text-guided drone geo-localization aims to identify a target region in a large-scale image gallery from a natural-language description. Existing methods mainly formulate this task as direct matching between an open-ended text query and candidate images. However, incomplete queries and highly similar candidates often make global cross-modal matching insufficient for reliable fine-grained localization. We propose UniGeo, a unified multimodal large language model (MLLM) for text-guided drone geo-localization. Built on a shared vision-language framework, UniGeo jointly supports geo-semantic understanding, cross-view semantic generation, and candidate-level verification. Specifically, it establishes stable correspondences among local scene elements, spatial relations, and language descriptions through geo-semantic learning, and further models semantic mappings between drone and satellite views through cross-view generation. Based on these capabilities, a plug-and-play verification module performs fine-grained discrimination among highly confusable candidates. We further introduce a multi-stage training strategy that progressively learns geo-semantic understanding, cross-view generation, and candidate verification, improving adaptation to text-guided geo-localization. Experiments demonstrate consistent improvements across multiple retrieval backbones. On GeoText-1652, UniGeo improves R@10 and mAP by 13.59 and 2.83 percentage points, respectively, validating its effectiveness for fine-grained text-guided drone geo-localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。