arXiv:2412.20597cs.CL2024-12中稿 · NoDaLiDa/Baltic-HL…

用GliNER提升爱沙尼亚语词形还原准确率,改进信息检索效果

GliLem: Leveraging GliNER for Contextualized Lemmatization in Estonian

  • 结合Vabamorf规则系统与GliNER外部消歧模块,实现上下文感知词形还原
  • 相比原系统提升10%准确率,在高k值下信息检索召回率稳定提高
  • 适合需要精准词形还原的爱沙尼亚语自然语言处理研究者

我们提出GliLem——一种针对爱沙尼亚语的新型混合词形还原系统,将高精度规则型形态分析器Vabamorf与基于GliNER(一个能以自然语言匹配文本片段与标签的开放词汇命名实体识别模型)的外部消歧模块相结合。利用预训练GliNER模型的灵活性,使Vabamorf的词形还原准确率相较原有消歧模块提升10%,并优于基于标记分类的基线方法。为评估词形还原精度提升对信息检索任务的影响,我们通过自动翻译英文DBpedia-Entity数据集,构建了首个爱沙尼亚语信息检索数据集。在该数据集上使用BM25算法对比多种标记归一化方法,发现词形还原显著优于简单词干提取;进一步提升词形消歧精度后,在高k值设置下,信息检索召回率获得小而稳定的提升。

原文摘要 · Abstract (English)

We present GliLem -- a novel hybrid lemmatization system for Estonian that enhances the highly accurate rule-based morphological analyzer Vabamorf with an external disambiguation module based on GliNER -- an open vocabulary NER model that is able to match text spans with text labels in natural language. We leverage the flexibility of a pre-trained GliNER model to improve the lemmatization accuracy of Vabamorf by 10% compared to its original disambiguation module and achieve an improvement over the token classification-based baseline. To measure the impact of improvements in lemmatization accuracy on the information retrieval downstream task, we first created an information retrieval dataset for Estonian by automatically translating the DBpedia-Entity dataset from English. We benchmark several token normalization approaches, including lemmatization, on the created dataset using the BM25 algorithm. We observe a substantial improvement in IR metrics when using lemmatization over simplistic stemming. The benefits of improving lemma disambiguation accuracy manifest in small but consistent improvement in the IR recall measure, especially in the setting of high k.

词形还原爱沙尼亚语NER信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。