arXiv:2505.10714cs.CL2025-05被引 4

测试大模型理解网格化地理数据的能力,发现视觉语言模型表现最佳。

GeoGrid-Bench: Can Foundation Models Understand Multimodal Gridded Geo-Spatial Data?

  • 构建涵盖16种气候变量的网格化地理数据基准
  • 3200个问答对覆盖单点到跨区域的复杂时空任务
  • 适合关注地理科学与大模型结合的研究者

我们提出GeoGrid-Bench,一个用于评估基础模型理解网格化地理空间数据能力的基准。地理空间数据因其密集的数值、强时空依赖性以及包含表格、热力图和地理可视化等多模态表示而具有独特挑战。为评估基础模型在该领域支持科学研究的潜力,该基准包含覆盖150个地点、长达多时段的16种气候变量的真实世界大规模数据,共生成约3200个问答对,基于8个领域专家设计的模板,反映科学家实际遇到的任务,从单点单时的简单查询到跨区域跨时段的复杂时空比较。评估表明,视觉-语言模型整体表现最佳,并提供了不同基础模型在各类地理空间任务中的优劣细粒度分析。该基准为大模型有效应用于地理空间数据分析及科研支持提供了更清晰的洞见。

原文摘要 · Abstract (English)

We present GeoGrid-Bench, a benchmark designed to evaluate the ability of foundation models to understand geo-spatial data in the grid structure. Geo-spatial datasets pose distinct challenges due to their dense numerical values, strong spatial and temporal dependencies, and unique multimodal representations including tabular data, heatmaps, and geographic visualizations. To assess how foundation models can support scientific research in this domain, GeoGrid-Bench features large-scale, real-world data covering 16 climate variables across 150 locations and extended time frames. The benchmark includes approximately 3,200 question-answer pairs, systematically generated from 8 domain expert-curated templates to reflect practical tasks encountered by human scientists. These range from basic queries at a single location and time to complex spatiotemporal comparisons across regions and periods. Our evaluation reveals that vision-language models perform best overall, and we provide a fine-grained analysis of the strengths and limitations of different foundation models in different geo-spatial tasks. This benchmark offers clearer insights into how foundation models can be effectively applied to geo-spatial data analysis and used to support scientific research.

地理信息多模态大模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。