首个土耳其语Text-to-SQL数据集,用于评估大模型跨语言能力
BIRDTurk: Adaptation of the BIRD Text-to-SQL Dataset to Turkish
- 通过受控翻译流程构建土耳其语Text-to-SQL数据集,保持原逻辑结构
- 人类评估准确率达98.15%,验证翻译质量可靠
- 揭示多语言模型在低资源语言中性能下降,适合跨语言研究者使用
Text-to-SQL系统在英语基准上表现强劲,但在形态丰富、资源匮乏的语言中行为仍不明确。我们提出BIRDTurk,首个土耳其语版BIRD基准数据集,通过受控翻译流程构建,适配数据库模式标识符至土耳其语,严格保留SQL查询与数据库的逻辑结构及执行语义。翻译质量基于中心极限定理确定样本量,确保95%置信度,人类评估准确率达98.15%。利用BIRDTurk,我们评估了基于推理的提示、代理式多阶段推理和监督微调。结果表明,土耳其语导致一致性能下降,源于语言结构差异与预训练数据不足;而代理式推理展现更强跨语言鲁棒性。标准多语言基线微调困难,但现代指令微调模型可有效扩展。BIRDTurk为真实数据库条件下跨语言Text-to-SQL评估提供可控测试平台。我们发布训练与开发集以支持后续研究。
原文摘要 · Abstract (English)
Text-to-SQL systems have achieved strong performance on English benchmarks, yet their behavior in morphologically rich, low-resource languages remains largely unexplored. We introduce BIRDTurk, the first Turkish adaptation of the BIRD benchmark, constructed through a controlled translation pipeline that adapts schema identifiers to Turkish while strictly preserving the logical structure and execution semantics of SQL queries and databases. Translation quality is validated on a sample size determined by the Central Limit Theorem to ensure 95% confidence, achieving 98.15% accuracy on human-evaluated samples. Using BIRDTurk, we evaluate inference-based prompting, agentic multi-stage reasoning, and supervised fine-tuning. Our results reveal that Turkish introduces consistent performance degradation, driven by both structural linguistic divergence and underrepresentation in LLM pretraining, while agentic reasoning demonstrates stronger cross-lingual robustness. Supervised fine-tuning remains challenging for standard multilingual baselines but scales effectively with modern instruction-tuned models. BIRDTurk provides a controlled testbed for cross-lingual Text-to-SQL evaluation under realistic database conditions. We release the training and development splits to support future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。