为罗曼什语五种方言和标准语设计的词形还原工具,准确率达95%。
RUMLEM: A Dictionary-Based Lemmatizer for Romansh
- 基于社区共建的形态数据库构建,覆盖五种方言及标准语。
- 在典型文本中可还原77%-84%的词汇,方言识别准确率达95%。
- 适用于方言识别与语言分类,适合小语种NLP研究者使用。
词形还原是许多自然语言处理应用中的关键步骤。本文提出RUMLEM,一个涵盖罗曼什语五种主要方言及超区域标准语鲁曼茨格里申(Rumantsch Grischun)的词形还原器。它基于为每种罗曼什语变体构建的全面、社区驱动的形态数据库,使RUMLEM能覆盖典型罗曼什文本中77%至84%的词汇。由于每种方言均有独立数据库,RUMLEM还可用于变体感知的语言分类。在3万篇长度各异的罗曼什语文本上评估显示,其在95%的情况下正确识别了语言变体。此外,概念验证表明,基于该词形还原器可实现罗曼什语与非罗曼什语的分类。
原文摘要 · Abstract (English)
Lemmatization -- the task of mapping an inflected word form to its dictionary form -- is a crucial component of many NLP applications. In this paper, we present RUMLEM, a lemmatizer that covers the five main varieties of Romansh as well as the supra-regional standard variety Rumantsch Grischun. It is based on comprehensive, community-driven morphological databases for Romansh, enabling RUMLEM to cover 77-84% of the words in a typical Romansh text. Since there is a dedicated database for each Romansh variety, an additional application of RUMLEM is variety-aware language classification. Evaluation on 30'000 Romansh texts of varying lengths shows that RUMLEM correctly identifies the variety in 95% of cases. In addition, a proof of concept demonstrates the feasibility of Romansh vs. non-Romansh language classification based on the lemmatizer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。