arXiv:2501.18750cs.CLcs.IR2025-01中稿 · NoDaLiDa/Baltic-HL…被引 2

改进投影式数据迁移,提升低资源语言的跨语言命名实体识别效果

Revisiting Projection-based Data Transfer for Cross-Lingual Named Entity Recognition in Low-Resource Languages

  • 用反向翻译优化词对齐,提高标注投影准确性
  • 提出新匹配机制,将源语言实体与目标候选精准对应
  • 在57种语言上验证,优于现有投影方法

跨语言命名实体识别通过语言间知识迁移来识别和分类命名实体,对低资源语言尤其重要。本文表明,基于数据的跨语言迁移方法在低资源语言场景下有效,甚至优于多语言语言模型。针对低资源语言的跨语言命名实体识别,本文对标注投影步骤提出两项关键改进:一是利用反向翻译优化词对齐以提升精度;二是提出一种新的形式化投影方法,用于匹配源语言实体与目标语言候选实体。在涵盖57种语言的两个数据集上进行大量实验,结果表明该方法在低资源设置下显著优于现有投影方法。研究凸显了投影式数据迁移在低资源跨语言命名实体识别中作为模型方法替代方案的稳健性。

原文摘要 · Abstract (English)

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to identify and classify named entities, making it particularly useful for low-resource languages. We show that the data-based cross-lingual transfer method is an effective technique for crosslingual NER and can outperform multilingual language models for low-resource languages. This paper introduces two key enhancements to the annotation projection step in cross-lingual NER for low-resource languages. First, we explore refining word alignments using back-translation to improve accuracy. Second, we present a novel formalized projection approach of matching source entities with extracted target candidates. Through extensive experiments on two datasets spanning 57 languages, we demonstrated that our approach surpasses existing projectionbased methods in low-resource settings. These findings highlight the robustness of projection-based data transfer as an alternative to model-based methods for crosslingual named entity recognition in lowresource languages.

跨语言命名实体低资源数据迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。