构建新分类体系与合成数据集,提升文本转SQL的真实性和多样性。
Beyond SELECT: A Comprehensive Taxonomy-Guided Benchmark for Real-World Text-to-SQL Translation
- 基于意图、语句类型等维度建立文本转SQL新分类体系。
- 合成数据集SQL-Synth覆盖更广,显著优于现有基准。
- 揭示大模型在真实场景下的短板,适合训练数据构建者参考。
文本转SQL数据集对模型训练与评估至关重要,但现有数据集覆盖有限,难以反映真实应用的多样性。为此,我们提出一个包含核心意图、语句类型、语法结构和关键操作维度的新分类体系。基于该体系,我们评估了Spider、Bird等常用公开数据集,发现其覆盖与多样性存在不足。随后,我们设计了一种基于分类体系的合成数据流水线,生成新数据集SQL-Synth。该方法结合分类体系与大语言模型(LLMs),确保数据集涵盖真实世界文本转SQL的广度与复杂性。大量分析与实验结果表明,该分类体系有效,SQL-Synth相比现有基准具有更高多样性与覆盖度。同时,我们发现现有大模型通常未能充分捕捉全部场景,导致在SQL-Synth上表现受限,但微调可显著提升性能。该分类体系具有重要影响,不仅支持数据集与模型性能的全面分析,还可指导大模型训练数据的构建。
原文摘要 · Abstract (English)
Text-to-SQL datasets are essential for training and evaluating text-to-SQL models, but existing datasets often suffer from limited coverage and fail to capture the diversity of real-world applications. To address this, we propose a novel taxonomy for text-to-SQL classification based on dimensions including core intents, statement types, syntax structures, and key actions. Using this taxonomy, we evaluate widely used public text-to-SQL datasets (e.g., Spider and Bird) and reveal limitations in their coverage and diversity. We then introduce a taxonomy-guided dataset synthesis pipeline, yielding a new dataset named SQL-Synth. This approach combines the taxonomy with Large Language Models (LLMs) to ensure the dataset reflects the breadth and complexity of real-world text-to-SQL applications. Extensive analysis and experimental results validate the effectiveness of our taxonomy, as SQL-Synth exhibits greater diversity and coverage compared to existing benchmarks. Moreover, we uncover that existing LLMs typically fall short in adequately capturing the full range of scenarios, resulting in limited performance on SQL-Synth. However, fine-tuning can substantially improve their performance in these scenarios. The proposed taxonomy has significant potential impact, as it not only enables comprehensive analysis of datasets and the performance of different LLMs, but also guides the construction of training data for LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。