arXiv:2508.10667cs.CVcs.AI2025-08被引 4

用卫星图与街景图对齐,提升大模型精准定位地址能力

AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models

  • 引入卫星图作为宏观线索,与街景图进行跨视图对齐
  • 在匹兹堡和旧金山数据集上准确率提升超9%和12%
  • 适合需要高精度城市地址定位的应用场景

大型视觉语言模型(LVLM)在国家或城市级别的粗粒度地理定位上表现优异,但在城市内部的细粒度街景定位上表现不佳。本文探索将全市范围内的地址定位能力融入LVLM,以实现基于街景图像的灵活地址相关问答。核心挑战在于街景视觉问答数据仅提供微观视觉线索,导致微调模型性能不足。为此,我们引入视角不变的卫星图像作为宏观线索,提出跨视图对齐调优方法,包括卫星图与街景图拼接机制及自动标签生成机制,通过跨视图匹配增强LVLM对街道分布的全局理解。提出的AddressVLM采用两阶段训练:跨视图对齐调优与地址定位调优。此外,我们基于匹兹堡和旧金山的图像地址定位数据集构建了两个街景视觉问答数据集。定性和定量评估表明,AddressVLM在两个数据集上的平均地址定位准确率分别比基线模型高出9%和12%。

原文摘要 · Abstract (English)

Large visual language models (LVLMs) have demonstrated impressive performance in coarse-grained geo-localization at the country or city level, but they struggle with fine-grained street-level localization within urban areas. In this paper, we explore integrating city-wide address localization capabilities into LVLMs, facilitating flexible address-related question answering using street-view images. A key challenge is that the street-view visual question-and-answer (VQA) data provides only microscopic visual cues, leading to subpar performance in fine-tuned models. To tackle this issue, we incorporate perspective-invariant satellite images as macro cues and propose cross-view alignment tuning including a satellite-view and street-view image grafting mechanism, along with an automatic label generation mechanism. Then LVLM's global understanding of street distribution is enhanced through cross-view matching. Our proposed model, named AddressVLM, consists of two-stage training protocols: cross-view alignment tuning and address localization tuning. Furthermore, we have constructed two street-view VQA datasets based on image address localization datasets from Pittsburgh and San Francisco. Qualitative and quantitative evaluations demonstrate that AddressVLM outperforms counterpart LVLMs by over 9% and 12% in average address localization accuracy on these two datasets, respectively.

地址定位跨视图对齐视觉语言模型街景图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。