构建企业级长文本语义对齐数据集,评估大模型在私有业务知识下的文本转SQL能力。
EntSQL: A Benchmark for Grounding Text-to-SQL in Long-Context Enterprise Knowledge

- 基于企业私有文档构建中英双语语义对齐数据集
- 最佳模型在长文档支持下仅达15.9%准确率,凸显知识依赖挑战
- 适用于评估大模型在真实企业场景中的上下文理解与推理能力
文本转SQL使用户可通过自然语言访问数据库,近年大模型显著提升了其性能。现有基准如Spider、BIRD和Spider~2.0侧重模式泛化、大规模数据库及真实工作流评估,但普遍忽视企业场景中依赖内部业务知识(如指标定义、报表规范、组织规则)的SQL生成需求。我们提出EntSQL,一个面向企业的长上下文文本转SQL评估基准,用于衡量大模型在私有业务文档中的语义定位能力。EntSQL包含跨五个业务领域的1,066个中英双语语义对齐样本,多数案例需超越问题与模式之外的领域知识,且涉及复杂SQL结构。在英文输入下,最优系统在提供长文档时仅达到15.9%的准确率,凸显了企业知识接地的难度。
原文摘要 · Abstract (English)
Text-to-SQL enables natural language access to databases, and recent LLMs have substantially advanced its capabilities. Existing benchmarks such as Spider, BIRD, and Spider~2.0 evaluate schema generalization, large-scale databases, and realistic workflows, but largely overlook enterprise scenarios where SQL generation depends on private business knowledge, such as internal metrics, reporting conventions, and organizational rules. We introduce EntSQL, an enterprise-oriented Text-to-SQL benchmark for evaluating long-context grounding over proprietary business documents. EntSQL contains 1,066 aligned Chinese-English semantic examples across five business domains, with most examples requiring domain knowledge beyond the question and schema and involving complex SQL structures. On English inputs, the best evaluated system reaches only 15.9\% when long-form documents are provided, highlighting the difficulty of grounding SQL generation in enterprise knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。