用GAN生成低资源语言数据,提升机器翻译效果
Generative-Adversarial Networks for Low-Resource Language Data Augmentation in Machine Translation
- 设计GAN模型生成低资源语言的单语数据
- 在2万句以下数据下仍能生成合理句子
- 适合低资源语言翻译研究者参考
神经机器翻译系统在低资源语言上表现不佳,因缺乏大规模语料库。人工数据标注成本高,我们提出使用生成对抗网络(GAN)进行数据增强。在模拟低资源环境下,仅用不到2万句训练数据,模型成功生成如“ask me that healthy lunch im cooking up”和“my grandfather work harder than your grandfather before”等合理句子。该方法首次探索了GAN在低资源机器翻译中的潜力,结果表明其具有进一步拓展应用的价值。
原文摘要 · Abstract (English)
Neural Machine Translation (NMT) systems struggle when translating to and from low-resource languages, which lack large-scale data corpora for models to use for training. As manual data curation is expensive and time-consuming, we propose utilizing a generative-adversarial network (GAN) to augment low-resource language data. When training on a very small amount of language data (under 20,000 sentences) in a simulated low-resource setting, our model shows potential at data augmentation, generating monolingual language data with sentences such as "ask me that healthy lunch im cooking up," and "my grandfather work harder than your grandfather before." Our novel data augmentation approach takes the first step in investigating the capability of GANs in low-resource NMT, and our results suggest that there is promise for future extension of GANs to low-resource NMT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。