首个阿拉伯方言文本转SQL数据集,聚焦摩洛哥方言。
Dialect2SQL: A Novel Text-to-SQL Dataset for Arabic Dialects with a Focus on Moroccan Darija
- 构建跨领域阿拉伯方言文本转SQL数据集,含9428对样本。
- 覆盖69个数据库,包含长模式、脏数据等真实挑战。
- 推动低资源语言NLP发展,适合方言与数据库研究者。
将自然语言问题(NLQ)转换为可执行SQL查询的文本转SQL任务近年来受到广泛关注,使非技术人员能与关系型数据库交互。尽管已有如SPIDER和WikiSQL等基准数据集推动模型发展,其他如SEDE和BIRD也引入了更复杂的现实场景挑战,但这些数据集主要针对英语、中文等高资源语言。本文提出Dialect2SQL,首个大规模、跨领域的阿拉伯语方言文本转SQL数据集,涵盖摩洛哥方言。数据集包含69个不同领域的数据库,共9,428组自然语言问题与对应SQL查询对。除长模式、脏值、复杂查询等常见挑战外,还融入摩洛哥方言特有的复杂性——多种语言来源、大量借词及独特表达方式。该数据集对文本转SQL社区及低资源语言资源建设具有重要价值。
原文摘要 · Abstract (English)
The task of converting natural language questions (NLQs) into executable SQL queries, known as text-to-SQL, has gained significant interest in recent years, as it enables non-technical users to interact with relational databases. Many benchmarks, such as SPIDER and WikiSQL, have contributed to the development of new models and the evaluation of their performance. In addition, other datasets, like SEDE and BIRD, have introduced more challenges and complexities to better map real-world scenarios. However, these datasets primarily focus on high-resource languages such as English and Chinese. In this work, we introduce Dialect2SQL, the first large-scale, cross-domain text-to-SQL dataset in an Arabic dialect. It consists of 9,428 NLQ-SQL pairs across 69 databases in various domains. Along with SQL-related challenges such as long schemas, dirty values, and complex queries, our dataset also incorporates the complexities of the Moroccan dialect, which is known for its diverse source languages, numerous borrowed words, and unique expressions. This demonstrates that our dataset will be a valuable contribution to both the text-to-SQL community and the development of resources for low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。