arXiv:2512.12084cs.IR2025-12AAAI被引 3

构建面向洪水管理的多表空间文本转SQL基准,提升真实场景下模型表现评估能力。

FloodSQL-Bench: A Retrieval-Augmented Benchmark for Geospatially-Grounded Text-to-SQL

  • 融合社会、基础设施与灾害数据,通过键值、空间和混合连接构建多表地理数据集
  • 在不同难度层级上系统评估大模型性能,揭示现有方法在复杂查询中的不足
  • 适用于灾备系统研发者、地理信息科学家及自然语言处理研究者

现有文本转SQL基准主要关注通用领域中的单表查询或有限连接,无法反映特定领域中多表及空间推理的复杂性。为弥补这一缺陷,我们提出FLOODSQL-BENCH,一个面向洪水管理领域的地理空间基准,通过键连接、空间连接及混合连接整合异构数据集。该基准结合社会、基础设施与灾害数据层,捕捉真实的洪水相关信息需求。我们以统一的检索增强生成设置系统评估近期大语言模型,并在不同难度层级上测量其表现。通过提供基于真实灾害管理数据的开放基准,FLOODSQL-BENCH为高风险应用领域中的文本转SQL研究建立了实用测试平台。

原文摘要 · Abstract (English)

Existing Text-to-SQL benchmarks primarily focus on single-table queries or limited joins in general-purpose domains, and thus fail to reflect the complexity of domain-specific, multi-table and geospatial reasoning, To address this limitation, we introduce FLOODSQL-BENCH, a geospatially grounded benchmark for the flood management domain that integrates heterogeneous datasets through key-based, spatial, and hybrid joins. The benchmark captures realistic flood-related information needs by combining social, infrastructural, and hazard data layers. We systematically evaluate recent large language models with the same retrieval-augmented generation settings and measure their performance across difficulty tiers. By providing a unified, open benchmark grounded in real-world disaster management data, FLOODSQL-BENCH establishes a practical testbed for advancing Text-to-SQL research in high-stakes application domains.

文本转SQL地理空间灾备系统大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。