arXiv:2509.25264cs.DBcs.AI2025-09被引 7

首个面向地理空间SQL生成的自动化评估框架,填补了大模型在地理数据库领域评测空白。

GeoSQL-Eval: First Evaluation of LLMs on PostGIS-Based NL2GeoSQL Queries

  • 构建首个基于PostGIS的自然语言转地理SQL评估体系
  • 涵盖14,178个任务实例与340个空间函数,覆盖多维认知能力
  • 适合地理信息科学、空间数据库及大模型评测研究者使用

大型语言模型(LLMs)在通用数据库的自然语言转SQL(NL2SQL)任务中表现优异,但扩展至地理空间SQL(GeoSQL)时,因涉及空间数据类型、函数调用和坐标系统,生成与执行难度显著提升。现有基准主要针对通用SQL,缺乏对GeoSQL的系统性评估框架。为此,我们提出GeoSQL-Eval——首个端到端自动化评估框架,并配套发布GeoSQL-Bench基准数据集。该数据集包含三个任务类别:概念理解、语法级SQL生成与模式检索,共涵盖14,178个实例、340个PostGIS函数及82个主题数据库。GeoSQL-Eval基于韦伯深度认知模型(DOK),涵盖四个认知维度、五个能力层级和二十种任务类型,实现从知识获取、语法生成到语义对齐、执行准确率与鲁棒性的全流程评估。我们对24个代表性模型在六类任务中进行测评,并采用熵权法结合统计分析,揭示性能差异、常见错误模式与资源消耗特征。最终发布公开的GeoSQL-Eval排行榜平台,支持持续测试与全球对比。本工作拓展了NL2GeoSQL范式,提供标准化、可解释且可扩展的评估框架,为地理空间信息科学及相关应用提供重要参考。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown strong performance in natural language to SQL (NL2SQL) tasks within general databases. However, extending to GeoSQL introduces additional complexity from spatial data types, function invocation, and coordinate systems, which greatly increases generation and execution difficulty. Existing benchmarks mainly target general SQL, and a systematic evaluation framework for GeoSQL is still lacking. To fill this gap, we present GeoSQL-Eval, the first end-to-end automated evaluation framework for PostGIS query generation, together with GeoSQL-Bench, a benchmark for assessing LLM performance in NL2GeoSQL tasks. GeoSQL-Bench defines three task categories-conceptual understanding, syntax-level SQL generation, and schema retrieval-comprising 14,178 instances, 340 PostGIS functions, and 82 thematic databases. GeoSQL-Eval is grounded in Webb's Depth of Knowledge (DOK) model, covering four cognitive dimensions, five capability levels, and twenty task types to establish a comprehensive process from knowledge acquisition and syntax generation to semantic alignment, execution accuracy, and robustness. We evaluate 24 representative models across six categories and apply the entropy weight method with statistical analyses to uncover performance differences, common error patterns, and resource usage. Finally, we release a public GeoSQL-Eval leaderboard platform for continuous testing and global comparison. This work extends the NL2GeoSQL paradigm and provides a standardized, interpretable, and extensible framework for evaluating LLMs in spatial database contexts, offering valuable references for geospatial information science and related applications.

地理信息大模型评测空间数据库自然语言查询

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。