arXiv:2501.18197cs.LGcs.DB2025-01被引 10

揭示Text2SQL评估的两大根本缺陷,提醒别迷信现有指标。

Fundamental Challenges in Evaluating Text2SQL Solutions and Detecting Their Limitations

  • 提出统一分类框架,系统梳理Text2SQL的错误成因
  • 发现评测数据质量差与匹配函数偏差是两大核心问题
  • 适合研究评测方法、模型可解释性的从业者参考

本文深入探讨Text2SQL解决方案评估中的根本挑战,指出现有基准中聚合指标的潜在风险。识别出开放基准中两个未被充分关注的局限:(1) 评测数据存在质量问题,主要源于未能捕捉自然语言到结构化查询转换的不确定性(如自然语言歧义);(2) 使用不同匹配函数作为SQL等价性近似引入了偏差。为厘清这两点,我们提出一个涵盖所有Text2SQL局限的统一分类体系,并通过主流Text2SQL模型和基准的实证调研加以验证。结合真实案例说明各类局限的成因,并为每类提出潜在缓解方案。最后指出,实施这些缓解策略或自动应用该分类体系仍面临诸多开放性挑战。

原文摘要 · Abstract (English)

In this work, we dive into the fundamental challenges of evaluating Text2SQL solutions and highlight potential failure causes and the potential risks of relying on aggregate metrics in existing benchmarks. We identify two largely unaddressed limitations in current open benchmarks: (1) data quality issues in the evaluation data, mainly attributed to the lack of capturing the probabilistic nature of translating a natural language description into a structured query (e.g., NL ambiguity), and (2) the bias introduced by using different match functions as approximations for SQL equivalence. To put both limitations into context, we propose a unified taxonomy of all Text2SQL limitations that can lead to both prediction and evaluation errors. We then motivate the taxonomy by providing a survey of Text2SQL limitations using state-of-the-art Text2SQL solutions and benchmarks. We describe the causes of limitations with real-world examples and propose potential mitigation solutions for each category in the taxonomy. We conclude by highlighting the open challenges encountered when deploying such mitigation strategies or attempting to automatically apply the taxonomy.

Text2SQL评估挑战数据质量偏差分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。