arXiv:2502.13520cs.CL2025-02ACL被引 27

构建了覆盖19个阅读等级的阿拉伯语可读性评估数据集

A Large and Balanced Corpus for Fine-grained Arabic Readability Assessment

  • 构建69,441句、百万词级阿拉伯语可读性数据集,覆盖从幼儿园到研究生水平
  • 人工标注达成81.8%的标注者间一致性(加权卡帕值)
  • 开源数据集与评测基准,适合语言评估与教育研究

本文提出平衡阿拉伯语可读性评估语料库(BAREC),一个大规模、细粒度的阿拉伯语可读性评估数据集。BAREC包含69,441个句子,总计超过100万词,精心设计以覆盖从幼儿园到研究生阶段的19个阅读水平。语料库在体裁多样性、主题覆盖面和目标读者方面保持均衡,为评估阿拉伯语文本复杂度提供了全面资源。所有文本均由大型标注团队进行人工标注,平均两两标注者一致性(加权卡帕值)达81.8%,表明高度的一致性。除发布语料库外,本文还在不同细粒度级别上对自动可读性评估方法进行了基准测试,比较多种技术。结果揭示了阿拉伯语可读性建模中的挑战与机遇,展示了各类方法的竞争力。为支持研究与教育,BAREC已公开发布,附带详细标注指南与基准结果。

原文摘要 · Abstract (English)

This paper introduces the Balanced Arabic Readability Evaluation Corpus (BAREC), a large-scale, fine-grained dataset for Arabic readability assessment. BAREC consists of 69,441 sentences spanning 1+ million words, carefully curated to cover 19 readability levels, from kindergarten to postgraduate comprehension. The corpus balances genre diversity, topical coverage, and target audiences, offering a comprehensive resource for evaluating Arabic text complexity. The corpus was fully manually annotated by a large team of annotators. The average pairwise inter-annotator agreement, measured by Quadratic Weighted Kappa, is 81.8%, reflecting a high level of substantial agreement. Beyond presenting the corpus, we benchmark automatic readability assessment across different granularity levels, comparing a range of techniques. Our results highlight the challenges and opportunities in Arabic readability modeling, demonstrating competitive performance across various methods. To support research and education, we make BAREC openly available, along with detailed annotation guidelines and benchmark results.

可读性评估阿拉伯语数据集自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。