arXiv:2512.24235cs.CL2025-12Conference of the …被引 4

构建首个大规模阿拉伯语作文评分数据集,支持多维度自动评分研究

LAILA: A Large Trait-Based Dataset for Arabic Automated Essay Scoring

  • 收集7859篇阿拉伯语作文,标注七维评分:相关性、结构、词汇等
  • 提供跨提示与特定提示场景的基准测试结果,验证模型性能
  • 填补阿拉伯语自动作文评分数据空白,适合教育AI研究者使用

近年来,自动作文评分(AES)受到越来越多关注,但阿拉伯语领域的研究仍受限于缺乏公开数据集。为此,我们推出了目前最大的公开阿拉伯语自动作文评分数据集LAILA,包含7,859篇作文,涵盖七个维度的全局与特质评分:相关性、结构、词汇、风格、内容展开、语法和标点。本文详细阐述了数据集的设计、采集与标注流程,并在特定提示与跨提示场景下,使用最先进的阿拉伯语与英语模型进行了基准测试。LAILA填补了阿拉伯语自动作文评分研究的关键空白,为构建稳健的评分系统提供了重要支持。

原文摘要 · Abstract (English)

Automated Essay Scoring (AES) has gained increasing attention in recent years, yet research on Arabic AES remains limited due to the lack of publicly available datasets. To address this, we introduce LAILA, the largest publicly available Arabic AES dataset to date, comprising 7,859 essays annotated with holistic and trait-specific scores on seven dimensions: relevance, organization, vocabulary, style, development, mechanics, and grammar. We detail the dataset design, collection, and annotations, and provide benchmark results using state-of-the-art Arabic and English models in prompt-specific and cross-prompt settings. LAILA fills a critical need in Arabic AES research, supporting the development of robust scoring systems.

自动评分阿拉伯语教育AI数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。