为爱尔兰语构建分词资源并评估分词与形态边界对齐程度
MoirfEolas and Cr\'iochScore: Developing Resources for and the Evaluation of Tokenization Alignment with Irish Morphology
- 构建3.5万词的爱尔兰语形态标注数据集MoirfEolas
- Unigram模型在形态对齐上表现最优,但存在压缩与词表效率权衡
- 适合研究低资源语言形态分析或分词优化的学者使用
本文提出针对爱尔兰语的新分词资源及评估指标。构建了包含超过3.5万词的MoirfEolas数据集,每词标注其鼻化、前缀和后缀等形态成分;并提出CríochScore评估指标,用于衡量分词结果与形态边界的一致性。通过该指标评估常见分词算法,发现Unigram语言模型在形态对齐方面优于其他方法。同时揭示分词在形态对齐、压缩率与词表效率之间的权衡关系,为爱尔兰语自然语言处理提供实用指导。该数据集有助于缓解爱尔兰语作为低资源语言的困境;文中提出的构建方法可推广至其他语言,以建立专用形态资源。
原文摘要 · Abstract (English)
This paper presents new tokenization resources for Irish and evaluation measures of alignment with the morphological boundaries of the language. We present MoirfEolas, a dataset of over 35,000 Irish words mapped to their respective eclipses, prefixes and suffixes as well as an evaluation metric Cr\'iochScore, that evaluates the alignment of tokenizations with the morphological boundaries present in MoirfEolas. We evaluate common tokenization algorithms using Cr\'iochScore as well as intrinsic metrics present in the tokenization literature. We find that the Unigram Language Model aligns with Irish morphology more often than the other algorithms evaluated. We also find trade-offs between morphological-alignment of tokenization with both compression as well as vocabulary efficiency, providing practical insights for Irish natural language processing development. This dataset contributes towards combating the Irish language's low-resource status; moreover, the construction process reported in this paper can be emulated by other languages to create specialised morphological resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。