构建首个几内亚比绍克里奥尔语翻译数据集,揭示小规模目标域数据对低资源语言翻译的关键作用
Limitations of Religious Data and the Importance of the Target Domain: Towards Machine Translation for Guinea-Bissau Creole
- 用4万条宗教文本构建克里奥尔语翻译数据集,探索从宗教领域向通用领域迁移的方法
- 仅添加300条目标域句子,翻译性能显著提升,证明小规模数据收集的重要性
- 葡萄牙语→克里奥尔语翻译效果更优,与词汇重叠和形态复杂性有关,适合克里奥尔语研究者
我们引入了一个新的几内亚比绍克里奥尔语(Kiriol)机器翻译数据集,包含约4万对平行语句,对应英语和葡萄牙语。该数据集主要由宗教文本(《圣经》及耶和华见证人资料)构成,辅以少量通用领域数据(来自词典)。这反映了众多低资源语言的典型资源状况。我们训练了多种基于Transformer的模型,研究如何将宗教领域数据有效迁移到通用领域。结果发现,即使在训练中加入300条目标领域句子,也能显著提升翻译性能,凸显了为低资源语言开展小规模数据采集的必要性。此外,我们发现葡萄牙语到克里奥尔语的翻译模型平均表现优于其他语言对,并探究其与语言形态复杂度及克里奥尔语与词源语言之间词汇重叠程度的关系。整体上,我们希望本工作能推动克里奥尔语及相关机器翻译研究的发展。
原文摘要 · Abstract (English)
We introduce a new dataset for machine translation of Guinea-Bissau Creole (Kiriol), comprising around 40 thousand parallel sentences to English and Portuguese. This dataset is made up of predominantly religious data (from the Bible and texts from the Jehovah's Witnesses), but also a small amount of general domain data (from a dictionary). This mirrors the typical resource availability of many low resource languages. We train a number of transformer-based models to investigate how to improve domain transfer from religious data to a more general domain. We find that adding even 300 sentences from the target domain when training substantially improves the translation performance, highlighting the importance and need for data collection for low-resource languages, even on a small-scale. We additionally find that Portuguese-to-Kiriol translation models perform better on average than other source and target language pairs, and investigate how this relates to the morphological complexity of the languages involved and the degree of lexical overlap between creoles and lexifiers. Overall, we hope our work will stimulate research into Kiriol and into how machine translation might better support creole languages in general.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。