构建首个综合地理任务评测基准,检验大模型在地理空间与时间理解上的真实能力。
GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks
- 基于12个公开地理数据集,覆盖多样任务与领域
- 发现模型推理能力与规模对性能影响显著
- 适合评估地理认知类大模型,推动跨域泛化研究
在地理数据背景下,现有大型语言模型多在同质环境下被研究,严重限制了对其泛化能力的洞察。本文提出GeoBenchLLM,一个全面的基准,用于探测大模型在地理相关任务中的表现。我们精心选取了来自不同地理任务和领域的12个公开数据集,利用该基准评估了一系列大模型在地理空间与时间理解方面的能力。结果表明,模型的推理能力与参数规模对整体性能具有显著影响。GeoBenchLLM已开源,可访问https://github.com/Rfr2003/GeoBenchLLM。
原文摘要 · Abstract (English)
In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities. In this paper, we present \benchName, a comprehensive benchmark for probing LLMs on geo-related tasks. We leverage a careful selection of twelve publicly available datasets from diverse geo-related tasks and domains, and evaluate a set of LLMs on geo-spatial and temporal understanding using our benchmark. Our results show that reasoning and size have a strong impact on overall performance. GeoBenchLLM is publicly available at https://github.com/Rfr2003/GeoBenchLLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。