构建面向商业智能的NL2SQL评估基准,解决现有数据集不贴合实际业务问题的缺陷。
BIS: NL2SQL Service Evaluation Benchmark for Business Intelligence Scenarios
- 针对工业级BI场景设计新基准,涵盖典型业务提问类型。
- 提出两项新型语义相似度评估指标,更适配实际应用需求。
- 填补生产环境NL2SQL评估空白,适合企业研发与服务评测使用。
近年来,自然语言转结构化查询语言(NL2SQL)在商业智能(BI)应用中得到广泛应用。然而,现有NL2SQL基准不适合生产级BI场景,因其未覆盖常见的业务询问。为弥补这一缺口,我们开发了一个聚焦典型工业级BI问题的新基准。本文讨论了构建此类基准所面临的挑战,以及现有基准的不足。此外,我们在基准中引入反映常见业务查询的问题类别。最后,我们提出了两种新颖的语义相似度评估指标,用于衡量NL2SQL在BI应用和服务中的能力。
原文摘要 · Abstract (English)
NL2SQL (Natural Language to Structured Query Language) transformation has seen wide adoption in Business Intelligence (BI) applications in recent years. However, existing NL2SQL benchmarks are not suitable for production BI scenarios, as they are not designed for common business intelligence questions. To address this gap, we have developed a new benchmark focused on typical NL questions in industrial BI scenarios. We discuss the challenges of constructing a BI-focused benchmark and the shortcomings of existing benchmarks. Additionally, we introduce question categories in our benchmark that reflect common BI inquiries. Lastly, we propose two novel semantic similarity evaluation metrics for assessing NL2SQL capabilities in BI applications and services.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。