arXiv:2503.07214cs.CLcs.AI2025-03被引 2

用音标对比学习提升低资源语言零样本命名实体识别

Cross-Lingual IPA Contrastive Learning for Zero-Shot NER

  • 基于音标表示的跨语言对比学习,缩小语音差异
  • 在10种高资源语言对上实现平均性能显著提升
  • 适合研究低资源语言、跨语言迁移与音标表征的学者

现有零样本命名实体识别方法多依赖机器翻译,近期研究转向音素表示。本文探讨如何通过减少具有相似语音特征的语言间国际音标(IPA)转录差异,使高资源语言训练的模型有效应用于低资源语言。为此,我们构建了包含10组英语及高资源语言IPA对的CONLIPA数据集,覆盖10个常用语系。同时提出跨语言IPA对比学习方法IPAC。实验表明,该方法相较最优基线实现显著平均提升。

原文摘要 · Abstract (English)

Existing approaches to zero-shot Named Entity Recognition (NER) for low-resource languages have primarily relied on machine translation, whereas more recent methods have shifted focus to phonemic representation. Building upon this, we investigate how reducing the phonemic representation gap in IPA transcription between languages with similar phonetic characteristics enables models trained on high-resource languages to perform effectively on low-resource languages. In this work, we propose CONtrastive Learning with IPA (CONLIPA) dataset containing 10 English and high resource languages IPA pairs from 10 frequently used language families. We also propose a cross-lingual IPA Contrastive learning method (IPAC) using the CONLIPA dataset. Furthermore, our proposed dataset and methodology demonstrate a substantial average gain when compared to the best performing baseline.

命名实体识别跨语言音标表征零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。