arXiv:2601.14518cs.CL2026-01

生成符合业务逻辑的文本转SQL数据,提升企业级智能查询评估真实性

Business Logic-Driven Text-to-SQL Data Synthesis for Business Intelligence

  • 基于业务角色与流程生成真实场景数据
  • 98.44%业务真实性,显著优于现有方法
  • 揭示主流模型在复杂业务查询上仅42.86%准确率

在私有商业智能(BI)环境中评估文本转SQL代理面临真实领域数据稀缺的挑战。虽然合成数据可提供可扩展性,但现有生成方法无法捕捉业务真实性——即问题是否反映真实业务逻辑与工作流程。我们提出一种业务逻辑驱动的数据合成框架,生成基于业务角色、工作场景与流程的数据。同时,通过引入业务推理复杂度控制策略,提升数据质量,使问题需经历多样化的分析推理步骤。在生产规模的Salesforce数据库上的实验表明,所生成数据具有高业务真实性(98.44%),显著优于OmniSQL(+19.5%)和SQL-Factory(+54.7%),同时保持强问题-SQL对齐(98.59%)。合成数据还揭示,当前最先进文本转SQL模型在最复杂业务查询上仅达42.86%执行准确率。

原文摘要 · Abstract (English)

Evaluating Text-to-SQL agents in private business intelligence (BI) settings is challenging due to the scarcity of realistic, domain-specific data. While synthetic evaluation data offers a scalable solution, existing generation methods fail to capture business realism--whether questions reflect realistic business logic and workflows. We propose a Business Logic-Driven Data Synthesis framework that generates data grounded in business personas, work scenarios, and workflows. In addition, we improve the data quality by imposing a business reasoning complexity control strategy that diversifies the analytical reasoning steps required to answer the questions. Experiments on a production-scale Salesforce database show that our synthesized data achieves high business realism (98.44%), substantially outperforming OmniSQL (+19.5%) and SQL-Factory (+54.7%), while maintaining strong question-SQL alignment (98.59%). Our synthetic data also reveals that state-of-the-art Text-to-SQL models still have significant performance gaps, achieving only 42.86% execution accuracy on the most complex business queries.

文本转SQL数据合成业务智能评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。