用高资源语言数据提升低资源语言的文本标注效果
Cross-Lingual Transfer for Low-Resource Natural Language Processing
- 用多语言模型和翻译技术改进跨语言数据迁移
- 新方法在实体识别等任务上显著超越旧方法
- 适合关注多语言NLP与医疗文本处理的研究者
近年来,自然语言处理(NLP)取得显著进展,尤其得益于大型语言模型在诸多任务上实现前所未有的性能。然而,这些成果主要惠及英语等少数高资源语言,多数语言因训练数据和计算资源匮乏仍面临挑战。为此,本论文聚焦跨语言迁移学习,旨在利用高资源语言的数据与模型提升低资源语言的NLP表现。研究以序列标注任务为核心,包括命名实体识别、观点目标提取和论点挖掘。工作围绕三大目标展开:(1)通过改进翻译与标注投影技术推进基于数据的跨语言迁移;(2)利用最先进的多语言模型发展更优的模型迁移方法;(3)将方法应用于真实场景并开源资源以促进后续研究。具体提出T-Projection方法,借助文本到文本多语言模型与机器翻译系统,显著优于以往标注投影方法。针对模型迁移,引入一种受限解码算法,在零样本设置下提升文本到文本模型的跨语言序列标注能力。最后,构建了首个多语言文本到文本医学模型Medical mT5,展示了研究成果在实际应用中的价值。
原文摘要 · Abstract (English)
Natural Language Processing (NLP) has seen remarkable advances in recent years, particularly with the emergence of Large Language Models that have achieved unprecedented performance across many tasks. However, these developments have mainly benefited a small number of high-resource languages such as English. The majority of languages still face significant challenges due to the scarcity of training data and computational resources. To address this issue, this thesis focuses on cross-lingual transfer learning, a research area aimed at leveraging data and models from high-resource languages to improve NLP performance for low-resource languages. Specifically, we focus on Sequence Labeling tasks such as Named Entity Recognition, Opinion Target Extraction, and Argument Mining. The research is structured around three main objectives: (1) advancing data-based cross-lingual transfer learning methods through improved translation and annotation projection techniques, (2) developing enhanced model-based transfer learning approaches utilizing state-of-the-art multilingual models, and (3) applying these methods to real-world problems while creating open-source resources that facilitate future research in low-resource NLP. More specifically, this thesis presents a new method to improve data-based transfer with T-Projection, a state-of-the-art annotation projection method that leverages text-to-text multilingual models and machine translation systems. T-Projection significantly outperforms previous annotation projection methods by a wide margin. For model-based transfer, we introduce a constrained decoding algorithm that enhances cross-lingual Sequence Labeling in zero-shot settings using text-to-text models. Finally, we develop Medical mT5, the first multilingual text-to-text medical model, demonstrating the practical impact of our research on real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。