arXiv:2602.21480cs.DBcs.CL2026-02被引 1

评估大模型在复杂数据场景下的文本转SQL能力,发现传统指标失效。

Both Ends Count! Just How Good are LLM Agents at "Text-to-Big SQL"?

  • 提出面向大数据的新型评估指标,涵盖执行效率与成本
  • 实测显示GPT-4o速度提升12.16倍,而精度仅低7%
  • 适合关注生产级大模型数据库应用的研究者和工程师

文本转SQL与大数据分析各自有广泛基准测试,但联合评估研究有限。实际中,文本转SQL系统常嵌入大数据工作流,如大规模数据处理或交互式分析,我们称之为「文本转大库SQL」(Text-to-Big SQL)。然而现有文本转SQL基准测试范围狭窄,忽略规模化带来的成本与性能影响。例如,小数据集上微小的翻译错误,在大数据规模下会引发显著的成本与延迟开销,而传统指标完全未予考虑。本文通过引入新颖且具代表性的评估指标,解决这一被忽视的问题。研究聚焦于生产级大模型代理(LLM agents),该系统具备数据库无关性,可适应多样用户需求。通过对前沿模型的全面评估,发现传统文本转SQL指标不足以衡量大数据场景。相比之下,我们的新指标能准确反映执行效率、成本及数据规模的影响。例如,GPT-4o虽精度比顶尖后期模型低约7%,但速度提升最高达12.16倍;在大输入规模下,GPT-5.2的成本效益是Gemini 3 Pro的两倍以上。

原文摘要 · Abstract (English)

Text-to-SQL and Big Data are both extensively benchmarked fields, yet there is limited research that evaluates them jointly. In the real world, Text-to-SQL systems are often embedded with Big Data workflows, such as large-scale data processing or interactive data analytics. We refer to this as ``Text-to-Big SQL''. However, existing text-to-SQL benchmarks remain narrowly scoped and overlook the cost and performance implications that arise at scale. For instance, translation errors that are minor on small datasets lead to substantial cost and latency overheads as data scales, a relevant issue completely ignored by text-to-SQL metrics. In this paper, we overcome this overlooked challenge by introducing novel and representative metrics for evaluating Text-to-Big SQL. Our study focuses on production-level LLM agents, a database-agnostic system adaptable to diverse user needs. Via an extensive evaluation of frontier models, we show that text-to-SQL metrics are insufficient for Big Data. In contrast, our proposed text-to-Big SQL metrics accurately reflect execution efficiency, cost, and the impact of data scale. For example, GPT-4o compensates for roughly 7% lower accuracy than the top-performing later-generation models with up to a 12.16x speedup, while GPT-5.2 is more than twice as cost-effective as Gemini 3 Pro at large input scales.

大模型文本转SQL大数据评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。