arXiv:2509.01147cs.CL2025-09EMNLP被引 1

用实体对齐翻译提升非拉丁语种零样本命名实体识别

Zero-shot Cross-lingual NER via Mitigating Language Difference: An Entity-aligned Translation Perspective

  • 通过双语翻译对齐中日等非拉丁语种与英文实体
  • 在多语言维基百科上微调大模型增强跨语言对齐
  • 特别适用于中文、日文等结构差异大的低资源语言

跨语言命名实体识别(CL-NER)旨在将高资源语言的知识迁移到低资源语言。然而,现有零样本跨语言命名实体识别(ZCL-NER)方法主要针对拉丁字母语言(LSL),其共享语言特征有助于知识迁移。相比之下,对于非拉丁字母语言(NSL),如中文和日文,由于深层结构差异,性能常显著下降。为解决这一问题,我们提出一种实体对齐翻译(EAT)方法。该方法利用大语言模型(LLMs),采用双翻译策略对齐非拉丁语种与英语之间的实体。此外,我们在多语言维基百科数据上微调大模型,以增强源语言到目标语言的实体对齐能力。

原文摘要 · Abstract (English)

Cross-lingual Named Entity Recognition (CL-NER) aims to transfer knowledge from high-resource languages to low-resource languages. However, existing zero-shot CL-NER (ZCL-NER) approaches primarily focus on Latin script language (LSL), where shared linguistic features facilitate effective knowledge transfer. In contrast, for non-Latin script language (NSL), such as Chinese and Japanese, performance often degrades due to deep structural differences. To address these challenges, we propose an entity-aligned translation (EAT) approach. Leveraging large language models (LLMs), EAT employs a dual-translation strategy to align entities between NSL and English. In addition, we fine-tune LLMs using multilingual Wikipedia data to enhance the entity alignment from source to target languages.

跨语言命名实体识别大模型非拉丁语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。