arXiv:2503.16326cs.AI2025-03被引 17

OmniGeo让大模型读懂地理数据,提升空间智能任务表现

OmniGeo: Towards a Multimodal Large Language Models for Geospatial Artificial Intelligence

  • 专为地理空间设计的多模态大模型,融合图像、文本与地理元数据
  • 在零样本任务中超越专用模型和现有大模型,表现更优
  • 适合遥感、城市规划等地理智能领域研究者使用

多模态大语言模型(MLLM)的快速发展为人工智能开辟了新方向,可整合文本、图像及空间信息等多元大规模数据。本文探索将多模态大模型应用于地理空间人工智能(GeoAI),该领域利用空间数据解决地理语义、健康地理、城市地理、城市感知和遥感等领域的挑战。我们提出面向地理应用的多模态大模型 OmniGeo,能够处理和分析异构数据源,包括卫星影像、地理元数据和文本描述。通过结合自然语言理解与空间推理能力,该模型显著提升了指令遵循能力和 GeoAI 系统的准确性。实验表明,其在多种地理空间任务上优于专用模型和现有大模型,有效应对多模态特性,并在零样本任务中达到有竞争力的表现。代码将于发表后公开。

原文摘要 · Abstract (English)

The rapid advancement of multimodal large language models (LLMs) has opened new frontiers in artificial intelligence, enabling the integration of diverse large-scale data types such as text, images, and spatial information. In this paper, we explore the potential of multimodal LLMs (MLLM) for geospatial artificial intelligence (GeoAI), a field that leverages spatial data to address challenges in domains including Geospatial Semantics, Health Geography, Urban Geography, Urban Perception, and Remote Sensing. We propose a MLLM (OmniGeo) tailored to geospatial applications, capable of processing and analyzing heterogeneous data sources, including satellite imagery, geospatial metadata, and textual descriptions. By combining the strengths of natural language understanding and spatial reasoning, our model enhances the ability of instruction following and the accuracy of GeoAI systems. Results demonstrate that our model outperforms task-specific models and existing LLMs on diverse geospatial tasks, effectively addressing the multimodality nature while achieving competitive results on the zero-shot geospatial tasks. Our code will be released after publication.

地理智能多模态模型大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。