arXiv:2606.07172cs.CVcs.AI2026-06中稿 · ICML

用文字监督提升视觉语言模型的地理空间理解能力

Textual Supervision Enhances Geospatial Representations in Vision-Language Models

论文配图:Textual Supervision Enhances Geospatial Representations in Vision-Language Models
图 1 · 摘自论文原文
  • 通过文本监督增强模型对地理空间信息的捕捉
  • 语言模态显著提升图像定位准确率,尤其在可定位场景中
  • 适合关注地理智能、多模态学习的研究者

地理空间理解是机器学习系统在图像地理定位和空间推理等任务中的关键但尚未充分探索的维度。本文分析了三类模型家族——仅视觉架构(如ViT)、视觉语言模型(如CLIP)以及大规模多模态基础模型(如LLaVA、Qwen和Gemma)——所获得的地理空间表征。通过对包含人物、地标和日常物品的图像聚类进行评估,按可定位程度分组,揭示了空间精度上的系统性差距,并表明文本监督能有效促进地理空间表征的学习。研究结果表明,语言作为编码空间上下文的有效互补模态,多模态学习是推动地理空间人工智能发展的关键方向。

原文摘要 · Abstract (English)

Geospatial understanding is a critical yet underexplored dimension in the development of machine learning systems for tasks such as image geolocation and spatial reasoning. In this work, we analyze the geospatial representations acquired by three model families: vision-only architectures (e.g., ViT), vision-language models (e.g., CLIP), and large-scale multimodal foundation models (e.g., LLaVA, Qwen, and Gemma). By evaluating across image clusters, including people, landmarks, and everyday objects, grouped based on the degree of localizability, we reveal systematic gaps in spatial accuracy and show that textual supervision enhances the learning of geospatial representations. Our findings suggest the role of language as an effective complementary modality for encoding spatial context and multimodal learning as a key direction for advancing geospatial AI.

地理空间多模态文本监督视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。