为阿拉伯语语法纠错构建平衡合成数据,提升低资源语言模型性能
Towards the Development of Balanced Synthetic Data for Correcting Grammatical Errors in Arabic: An Approach Based on Error Tagging Model and Synthetic Data Generating Model
- 用DeBERTav3和ARETA工具标注26类错误,精准识别正确句中的语法缺陷
- 基于反向翻译生成3000万+错句对,在QALB-14上达79.36%新高F1分数
- 适合研究低资源语言纠错、合成数据生成与多标签分类的学者参考
合成数据生成被广泛认为是提升神经语法纠错(GEC)系统质量的有效途径。然而,现有方法常缺乏多样性或过于简单,难以生成人类常见的复杂语法错误,尤其在阿拉伯语等低资源语言中更为明显。本文提出误差标记模型与合成数据生成模型,构建大规模阿拉伯语语法纠错合成数据集。误差标记模型利用DeBERTav3将正确句子按26种错误类型进行多标签分类,通过阿拉伯语错误类型标注工具(ARETA)指导任务;合成数据生成模型基于反向翻译,将错误标签附加至由误差标记模型生成的正确句前,使用ARAT5模型生成错误句子。在QALB-14和QALB-15测试集上,误差标记模型达到94.42% F1,为当前最佳水平。经合成数据训练后,语法纠错模型在QALB-14测试集上实现79.36%的新状态最优F1分数。共生成30,219,310条合成句对。
原文摘要 · Abstract (English)
Synthetic data generation is widely recognized as a way to enhance the quality of neural grammatical error correction (GEC) systems. However, current approaches often lack diversity or are too simplistic to generate the wide range of grammatical errors made by humans, especially for low-resource languages such as Arabic. In this paper, we will develop the error tagging model and the synthetic data generation model to create a large synthetic dataset in Arabic for grammatical error correction. In the error tagging model, the correct sentence is categorized into multiple error types by using the DeBERTav3 model. Arabic Error Type Annotation tool (ARETA) is used to guide multi-label classification tasks in an error tagging model in which each sentence is classified into 26 error tags. The synthetic data generation model is a back-translation-based model that generates incorrect sentences by appending error tags before the correct sentence that was generated from the error tagging model using the ARAT5 model. In the QALB-14 and QALB-15 Test sets, the error tagging model achieved 94.42% F1, which is state-of-the-art in identifying error tags in clean sentences. As a result of our syntactic data training in grammatical error correction, we achieved a new state-of-the-art result of F1-Score: 79.36% in the QALB-14 Test set. We generate 30,219,310 synthetic sentence pairs by using a synthetic data generation model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。