用合成数据提升阿拉伯语作文评分模型效果
Enhancing Arabic Automated Essay Scoring with Synthetic Data and Error Injection
- 用大模型生成跨水平的阿拉伯语作文
- 构建3040篇带错误注入的标注作文集
- 适配阿拉伯语学习者评分与教学反馈
自动作文评分(AES)在评估语言学习者写作质量、减轻阅卷负担和提供实时反馈方面起着关键作用。然而,标注作文数据集的缺乏制约了阿拉伯语AES系统的发展。本文利用大语言模型(LLMs)和Transformer模型生成用于阿拉伯语AES的合成作文。我们通过提示一个LLM生成覆盖欧洲共同语言参考框架(CEFR)各能力水平的作文,并提出并比较了两种错误注入方法。最终构建了一个包含3,040篇带错误的标注作文的数据集。此外,我们开发了一个基于BERT的阿拉伯语AES系统,可校准至CEFR等级。实验结果表明,所生成的合成数据集能有效提升阿拉伯语AES性能。代码与数据已公开。
原文摘要 · Abstract (English)
Automated Essay Scoring (AES) plays a crucial role in assessing language learners' writing quality, reducing grading workload, and providing real-time feedback. The lack of annotated essay datasets inhibits the development of Arabic AES systems. This paper leverages Large Language Models (LLMs) and Transformer models to generate synthetic Arabic essays for AES. We prompt an LLM to generate essays across the Common European Framework of Reference (CEFR) proficiency levels and introduce and compare two approaches to error injection. We create a dataset of 3,040 annotated essays with errors injected using our two methods. Additionally, we develop a BERT-based Arabic AES system calibrated to CEFR levels. Our experimental results demonstrate the effectiveness of our synthetic dataset in improving Arabic AES performance. We make our code and data publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。