用圣经数据提升低资源语言翻译,简单方法反最有效
From Priest to Doctor: Domain Adaptation for Low-Resource Neural Machine Translation
- 仅用圣经语料+词典+目标域单语数据做领域自适应
- 最简方法DALI在低资源翻译中表现最优
- 适合研究小样本、宗教文本翻译的学者参考
世界上许多语言缺乏训练高性能通用神经机器翻译(NMT)模型所需的数据,更别提领域专用模型。目前可用的平行语料通常仅为少量宗教文本。因此,针对低资源语言的领域自适应(DA)成为当前NMT的关键挑战,但研究仍不充分。本文在真实场景下评估了多种来自低资源NMT和领域自适应的方法:目标是在高资源语言与低资源语言之间进行翻译,仅可访问:a) 平行圣经数据,b) 双语词典,c) 高资源语言中的目标领域单语语料库。结果表明,所测试方法的效果各异,其中最简单的方法DALI表现最佳。我们进一步对DALI进行了小规模人工评估,结果显示仍需深入探究如何有效实现低资源NMT的领域自适应。
原文摘要 · Abstract (English)
Many of the world's languages have insufficient data to train high-performing general neural machine translation (NMT) models, let alone domain-specific models, and often the only available parallel data are small amounts of religious texts. Hence, domain adaptation (DA) is a crucial issue faced by contemporary NMT and has, so far, been underexplored for low-resource languages. In this paper, we evaluate a set of methods from both low-resource NMT and DA in a realistic setting, in which we aim to translate between a high-resource and a low-resource language with access to only: a) parallel Bible data, b) a bilingual dictionary, and c) a monolingual target-domain corpus in the high-resource language. Our results show that the effectiveness of the tested methods varies, with the simplest one, DALI, being most effective. We follow up with a small human evaluation of DALI, which shows that there is still a need for more careful investigation of how to accomplish DA for low-resource NMT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。