用大模型自动评分阿拉伯语作文,提升教育评估效率
Automated Scoring of Arabic Text Using Large Language Models: A Literature Review

- 构建五维分类体系,系统梳理阿拉伯语文本评分方法
- 对比分析现有研究的模型、数据与评分效果
- 为阿拉伯语教育评估提供可复用的技术框架
在现代教育体系中,自动文本评分(ATS)通过无需人工干预的方式实现学习者作答的规模化与一致性评价。近年来,大型语言模型(LLM)和阿拉伯语专用数据集的普及推动了该领域的复兴。本文聚焦基于LLM的阿拉伯语文本自动评价,涵盖短答案评分(ASAG)与作文评分(AES)。我们提出一个包含五个维度的结构化分类体系:应用领域、反馈生成能力、部署的LLM架构、与能力参照框架的对齐程度、提示工程策略。基于此分类体系,对现有研究进行比较分析,考察其方法论、数据集、评估指标及性能表现。研究结果表明,需持续开展以教育学为基础的研究,以提升阿拉伯语学习者评估质量,助力阿拉伯语教学发展。
原文摘要 · Abstract (English)
In modern educational systems, Automatic Text Scoring (ATS) plays a central role by enabling scalable and consistent evaluation of learner responses without human intervention. Recently, the increased accessibility of LLMs and Arabic-specific datasets has sparked renewed interest in this area. In this work, we investigate LLM-Based approaches for the automated evaluation of Arabic texts, focusing on both short answer grading (ASAG) and essay scoring (AES). We further introduce a structured taxonomy comprising five dimensions: application domain, feedback generation capability, LLM architecture deployed, alignment with competency referential frameworks, and prompt engineering strategy. By applying this taxonomy, we conduct a comparative analysis of existing studies, examining their methodological approaches, datasets, evaluation metrics, and reported performance. The findings highlight the need for sustained and pedagogically grounded research efforts in Arabic ATS, given its significance for improving educational quality across Arabic-speaking communities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。