用辅助领域数据提升低资源语言的特定领域翻译效果
Exploiting Domain-Specific Parallel Data on Multilingual Language Models for Low-resource Language Translation
- 用其他领域的平行语料微调或继续预训练多语言模型
- 领域差异越大,翻译性能下降越明显
- 为低资源语言提供针对性翻译策略
基于多语言序列到序列语言模型(msLMs)的神经机器翻译(NMT)系统在平行数据量少、语言表征不足时表现不佳,限制了低资源语言(LRLs)的领域特定NMT能力。为此,可利用辅助领域的平行数据对msLM进行微调或进一步预训练。本文评估了这两种技术在特定领域低资源语言翻译中的有效性,并研究了领域差异对NMT性能的影响。实验结果表明,合理利用辅助平行数据可显著提升性能。论文还提出了若干针对低资源语言领域NMT建模的数据使用策略。
原文摘要 · Abstract (English)
Neural Machine Translation (NMT) systems built on multilingual sequence-to-sequence Language Models (msLMs) fail to deliver expected results when the amount of parallel data for a language, as well as the language's representation in the model are limited. This restricts the capabilities of domain-specific NMT systems for low-resource languages (LRLs). As a solution, parallel data from auxiliary domains can be used either to fine-tune or to further pre-train the msLM. We present an evaluation of the effectiveness of these two techniques in the context of domain-specific LRL-NMT. We also explore the impact of domain divergence on NMT model performance. We recommend several strategies for utilizing auxiliary parallel data in building domain-specific NMT models for LRLs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。