arXiv:2410.08674cs.CL2024-10中稿 · ACL被引 21

构建阿拉伯语细粒度可读性标注规范,支持从幼儿园到研究生的19级评估

Guidelines for Fine-grained Sentence-level Arabic Readability Annotation

  • 基于教育者迭代训练,制定19级阿拉伯语可读性标注指南
  • 19级分类下标注者间一致性达81.8%(加权肯德尔系数)
  • 适用于阿拉伯语教学、NLP可读性模型评估与教育技术研究

本文介绍平衡阿拉伯语可读性评估语料库(BAREC)的标注规范,该语料库是用于阿拉伯语细粒度句级可读性评估的大规模资源。BAREC包含69,441个句子(超过100万词),标注覆盖从幼儿园到研究生共19个级别。基于Taha/Arabi21框架,通过母语阿拉伯语教育者反复培训优化了标注指南。研究强调了影响可读性的语言、教学与认知因素,并报告了高标注一致性:在最后标注阶段,加权肯德尔相关系数达到81.8%(属实质性至优秀一致)。此外,我们在多种分类粒度(19级、7级、5级、3级)上对自动可读性模型进行了基准测试。该语料库和标注指南已公开发布。

原文摘要 · Abstract (English)

This paper presents the annotation guidelines of the Balanced Arabic Readability Evaluation Corpus (BAREC), a large-scale resource for fine-grained sentence-level readability assessment in Arabic. BAREC includes 69,441 sentences (1M+ words) labeled across 19 levels, from kindergarten to postgraduate. Based on the Taha/Arabi21 framework, the guidelines were refined through iterative training with native Arabic-speaking educators. We highlight key linguistic, pedagogical, and cognitive factors in determining readability and report high inter-annotator agreement: Quadratic Weighted Kappa 81.8% (substantial/excellent agreement) in the last annotation phase. We also benchmark automatic readability models across multiple classification granularities (19-, 7-, 5-, and 3-level). The corpus and guidelines are publicly available.

可读性评估阿拉伯语标注规范教育NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。