为印尼语语法纠错构建高效高质量语料库框架。
A Simple Yet Effective Corpus Construction Framework for Indonesian Grammatical Error Correction
- 基于可扩展框架自动生成印尼语纠错语料。
- 利用大模型提升标注效率,显著降低人工成本。
- 适用于低资源语言的语法纠错研究,开源可用。
目前,语法纠错(GEC)研究主要集中在英语、中文等通用语言上,许多低资源语言缺乏可访问的评估语料库。如何高效构建低资源语言的高质量评估语料库成为关键挑战。本文提出一种语料库构建框架,以印尼语为例,利用该框架构建了面向印尼语的GEC评估语料库,解决了现有语料库的局限性。此外,我们探究了使用GPT-3.5-Turbo和GPT-4等大语言模型在GEC标注中提升效率的可行性,结果表明大模型在低资源语言场景下具有显著潜力。代码与语料库已公开于https://github.com/GKLMIP/GEC-Construction-Framework。
原文摘要 · Abstract (English)
Currently, the majority of research in grammatical error correction (GEC) is concentrated on universal languages, such as English and Chinese. Many low-resource languages lack accessible evaluation corpora. How to efficiently construct high-quality evaluation corpora for GEC in low-resource languages has become a significant challenge. To fill these gaps, in this paper, we present a framework for constructing GEC corpora. Specifically, we focus on Indonesian as our research language and construct an evaluation corpus for Indonesian GEC using the proposed framework, addressing the limitations of existing evaluation corpora in Indonesian. Furthermore, we investigate the feasibility of utilizing existing large language models (LLMs), such as GPT-3.5-Turbo and GPT-4, to streamline corpus annotation efforts in GEC tasks. The results demonstrate significant potential for enhancing the performance of LLMs in low-resource language settings. Our code and corpus can be obtained from https://github.com/GKLMIP/GEC-Construction-Framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。