arXiv:2409.19014cs.CLcs.IR2024-09被引 10

用大模型模拟专家评估,解决文本转SQL评测中的误判问题。

FLEX: Expert-level False-Less EXecution Metric for Reliable Text-to-SQL Benchmark

  • 用LLM模拟人类专家,基于上下文和复杂标准评估查询
  • 与专家一致性提升至87.04(kappa),平均性能提高2.6点
  • 适合关注评测公正性与模型真实能力的研究者

文本转SQL系统在各行业日益重要,使非技术人员也能执行复杂数据操作。随着系统愈发复杂,准确评估方法的需求也日益迫切。然而,当前最常用的执行准确率(EX)仍存在大量误判。本文提出FLEX(False-Less EXecution),利用大语言模型(LLMs)模拟人类专家级的SQL查询评估,结合全面上下文与精细标准,将与人类专家的一致性从62提升至87.04(Cohen's kappa)。大规模实验揭示:(1) 模型在Spider和BIRD基准上的平均性能提升超2.6点,显著改变排名;(2) EX的低估主要源于标注质量缺陷;(3) 难题上模型表现常被高估。本工作推动了文本转SQL评估的准确性与细致性,可能重塑对当前顶尖性能的认知。

原文摘要 · Abstract (English)

Text-to-SQL systems have become crucial for translating natural language into SQL queries in various industries, enabling non-technical users to perform complex data operations. The need for accurate evaluation methods has increased as these systems have grown more sophisticated. However, the Execution Accuracy (EX), the most prevalent evaluation metric, still shows many false positives and negatives. Thus, this paper introduces FLEX (False-Less EXecution), a novel approach to evaluating text-to-SQL systems using large language models (LLMs) to emulate human expert-level evaluation of SQL queries. Our metric improves agreement with human experts (from 62 to 87.04 in Cohen's kappa) with comprehensive context and sophisticated criteria. Our extensive experiments yield several key insights: (1) Models' performance increases by over 2.6 points on average, substantially affecting rankings on Spider and BIRD benchmarks; (2) The underestimation of models in EX primarily stems from annotation quality issues; and (3) Model performance on particularly challenging questions tends to be overestimated. This work contributes to a more accurate and nuanced evaluation of text-to-SQL systems, potentially reshaping our understanding of state-of-the-art performance in this field.

文本转SQL评估指标大模型数据评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。