构建可解释的文本转SQL表示空间,发现并填补模型鲁棒性短板
SQLSpace: A Representation Space for Text-to-SQL to Discover and Mitigate Robustness Gaps
- 基于少量人工干预生成可解释的文本转SQL表示空间
- 揭示不同数据集在组合逻辑上的差异及模型性能隐藏模式
- 支持精准重写查询提升模型效果,适合评估与改进模型
我们提出 SQLSpace,一种人类可理解、可泛化且紧凑的文本转SQL示例表示空间,仅需极少人工干预即可生成。通过三个应用场景展示其有效性:(i) 精确对比主流文本转SQL基准的数据构成,识别各数据集评估的独特维度;(ii) 在整体准确率之外,深入分析模型在细粒度上的表现差异;(iii) 基于学习到的正确性估计,进行针对性查询重写以提升模型性能。结果表明,仅靠原始样本难以实现的分析成为可能:揭示了不同基准间的组合差异,暴露了被准确率掩盖的性能规律,并支持对查询成功概率的建模。
原文摘要 · Abstract (English)
We introduce SQLSpace, a human-interpretable, generalizable, compact representation for text-to-SQL examples derived with minimal human intervention. We demonstrate the utility of these representations in evaluation with three use cases: (i) closely comparing and contrasting the composition of popular text-to-SQL benchmarks to identify unique dimensions of examples they evaluate, (ii) understanding model performance at a granular level beyond overall accuracy scores, and (iii) improving model performance through targeted query rewriting based on learned correctness estimation. We show that SQLSpace enables analysis that would be difficult with raw examples alone: it reveals compositional differences between benchmarks, exposes performance patterns obscured by accuracy alone, and supports modeling of query success.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。