用图模型提升极低资源语言的翻译能力,效果远超传统方法。
BhashaSetu: Cross-Lingual Knowledge Transfer from High-Resource to Extreme Low-Resource Languages
- 基于图神经网络构建跨语言词表示,增强低资源语言理解。
- 在米佐语、卡西语上实现13%的词性标注准确率提升。
- 适合研究低资源语言处理或跨语言迁移的学者参考。
尽管自然语言处理取得显著进展,但为低资源语言开发有效系统仍面临巨大挑战,其性能通常远低于高资源语言,主要由于数据稀缺和语言资源不足。跨语言知识迁移成为解决该问题的有前景方法,通过利用高资源语言的资源来弥补低资源语言的短板。本文研究如何将高资源语言中的语言知识迁移到低资源语言,其中标注训练样本仅数百个。聚焦句子级与词级任务,提出一种新方法GETR(图增强词表示),并采用两种基线方法:隐藏层增强与通过词翻译传递词嵌入。实验表明,基于图神经网络的方法显著优于现有多种语言及跨语言基线模型,在真正低资源语言(米佐语、卡西语)的词性标注任务上分别提升13个百分点;在模拟低资源语言(马拉地语、孟加拉语、马拉雅拉姆语)上,情感分类与命名实体识别任务的宏平均F1值分别提升20和27个百分点。还对迁移机制进行了深入分析,识别出影响成功迁移的关键因素。
原文摘要 · Abstract (English)
Despite remarkable advances in natural language processing, developing effective systems for low-resource languages remains a formidable challenge, with performances typically lagging far behind high-resource counterparts due to data scarcity and insufficient linguistic resources. Cross-lingual knowledge transfer has emerged as a promising approach to address this challenge by leveraging resources from high-resource languages. In this paper, we investigate methods for transferring linguistic knowledge from high-resource languages to low-resource languages, where the number of labeled training instances is in hundreds. We focus on sentence-level and word-level tasks. We introduce a novel method, GETR (Graph-Enhanced Token Representation) for cross-lingual knowledge transfer along with two adopted baselines (a) augmentation in hidden layers and (b) token embedding transfer through token translation. Experimental results demonstrate that our GNN-based approach significantly outperforms existing multilingual and cross-lingual baseline methods, achieving 13 percentage point improvements on truly low-resource languages (Mizo, Khasi) for POS tagging, and 20 and 27 percentage point improvements in macro-F1 on simulated low-resource languages (Marathi, Bangla, Malayalam) across sentiment classification and NER tasks respectively. We also present a detailed analysis of the transfer mechanisms and identify key factors that contribute to successful knowledge transfer in this linguistic context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。