arXiv:2411.04588cs.CLcs.AI2024-11被引 8

用ChatGPT构建大规模阿拉伯语语法纠错语料库,解决数据稀缺问题。

Tibyan Corpus: Balanced and Comprehensive Error Coverage Corpus Using ChatGPT for Arabic Grammatical Error Correction

  • 基于书本原文生成带错句,用ChatGPT批量扩充数据。
  • 经语言专家验证,含49类错误,总词数约60万。
  • 适合研究阿拉伯语语法纠错与低资源语言NLP的学者。

自然语言处理常通过文本数据增强缓解样本量不足问题。本文以阿拉伯语为对象,针对其语法错误纠正(GEC)资源匮乏的问题,构建名为Tibyan的语料库。现有研究主要依赖QALB-14和QALB-15两个数据集,共约20,500对平行句,数量偏低。为此,本文利用ChatGPT,以阿拉伯语书籍中的正确句子(引导句)为基础,生成包含多种语法错误的对应句,实现数据扩充。通过多源文本收集与预处理,结合专家评审确保生成句质量,并反复迭代优化。最终使用阿拉伯语错误类型标注工具(ARETA)分析发现,该语料库涵盖49种错误类型,包括拼写、形态、句法、语义、标点、合并与拆分等七类。语料总量约为60万词符。

原文摘要 · Abstract (English)

Natural language processing (NLP) utilizes text data augmentation to overcome sample size constraints. Increasing the sample size is a natural and widely used strategy for alleviating these challenges. In this study, we chose Arabic to increase the sample size and correct grammatical errors. Arabic is considered one of the languages with limited resources for grammatical error correction (GEC). Furthermore, QALB-14 and QALB-15 are the only datasets used in most Arabic grammatical error correction research, with approximately 20,500 parallel examples, which is considered low compared with other languages. Therefore, this study aims to develop an Arabic corpus called "Tibyan" for grammatical error correction using ChatGPT. ChatGPT is used as a data augmenter tool based on a pair of Arabic sentences containing grammatical errors matched with a sentence free of errors extracted from Arabic books, called guide sentences. Multiple steps were involved in establishing our corpus, including the collection and pre-processing of a pair of Arabic texts from various sources, such as books and open-access corpora. We then used ChatGPT to generate a parallel corpus based on the text collected previously, as a guide for generating sentences with multiple types of errors. By engaging linguistic experts to review and validate the automatically generated sentences, we ensured that they were correct and error-free. The corpus was validated and refined iteratively based on feedback provided by linguistic experts to improve its accuracy. Finally, we used the Arabic Error Type Annotation tool (ARETA) to analyze the types of errors in the Tibyan corpus. Our corpus contained 49 of errors, including seven types: orthography, morphology, syntax, semantics, punctuation, merge, and split. The Tibyan corpus contains approximately 600 K tokens.

语法纠错阿拉伯语数据增强ChatGPT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。