arXiv:2506.18148cs.CLcs.AI2025-06被引 8

构建了7.7万词的古兰经形态标注语料库,支持语言学研究与跨资源关联。

QuranMorph: Morphologically Annotated Quranic Corpus

  • 由三位专家手动标注每个词的词元和词性,确保高精度。
  • 采用40类细粒度词性标签和200万词的词典资源进行标注。
  • 可与多种语言学资源互联,适合阿拉伯语研究者使用。

我们提出了QuranMorph语料库,这是针对古兰经(共77,429个词)的形态标注语料库。每个词均由三位专家手工进行词元化并标注词性。词元化基于Qabas词典数据库,该数据库链接110个词典与语料库,总规模达200万词。词性标注采用细粒度的SAMA/Qabas标注集,包含40个标签。本文表明,丰富的词元化和词性标注使QuranMorph能与多种语言学资源实现互连。该语料库为开源,作为SinaLab资源公开发布于https://sina.birzeit.edu/quran。

原文摘要 · Abstract (English)

We present the QuranMorph corpus, a morphologically annotated corpus for the Quran (77,429 tokens). Each token in the QuranMorph was manually lemmatized and tagged with its part-of-speech by three expert linguists. The lemmatization process utilized lemmas from Qabas, an Arabic lexicographic database linked with 110 lexicons and corpora of 2 million tokens. The part-of-speech tagging was performed using the fine-grained SAMA/Qabas tagset, which encompasses 40 tags. As shown in this paper, this rich lemmatization and POS tagset enabled the QuranMorph corpus to be inter-linked with many linguistic resources. The corpus is open-source and publicly available as part of the SinaLab resources at (https://sina.birzeit.edu/quran)

古兰经形态标注阿拉伯语语料库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。