用多语言模型提升低资源印欧语翻译质量,效果显著优于单语言模型。
Together We Can: Multilingual Automatic Post-Editing for Low-Resource Languages
- 利用语言相似性构建多语言自动后编辑模型,生成合成数据增强训练。
- 多任务学习使模型在低资源语言上提升2.5~2.39 TER点,最优达+1.44。
- 适合需要跨语言迁移与低资源翻译优化的研究者与实践者。
本探索性研究考察了多语言自动后编辑(APE)系统在提升低资源印欧语言机器翻译质量方面的潜力。聚焦英语-马拉地语和英语-印地语两个密切相关语言对,利用其语言相似性构建稳健的多语言APE模型。为促进跨语言迁移,生成合成的印地语-马拉地语和马拉地语-印地语APE三元组数据。同时引入质量评估(QE)与APE的多任务学习框架。实验表明,APE与QE具有互补性,且多任务学习有效支持领域自适应。结果证明,多语言APE模型在英语-印地语和英语-马拉地语任务上分别优于单语言模型2.5和2.39 TER点;通过多任务学习、数据增强及领域适配,进一步提升1.29、1.44、0.53、0.45和0.35、0.45 TER点。相关合成数据、代码与模型已公开于https://github.com/cfiltnlp/Multilingual-APE。
原文摘要 · Abstract (English)
This exploratory study investigates the potential of multilingual Automatic Post-Editing (APE) systems to enhance the quality of machine translations for low-resource Indo-Aryan languages. Focusing on two closely related language pairs, English-Marathi and English-Hindi, we exploit the linguistic similarities to develop a robust multilingual APE model. To facilitate cross-linguistic transfer, we generate synthetic Hindi-Marathi and Marathi-Hindi APE triplets. Additionally, we incorporate a Quality Estimation (QE)-APE multi-task learning framework. While the experimental results underline the complementary nature of APE and QE, we also observe that QE-APE multitask learning facilitates effective domain adaptation. Our experiments demonstrate that the multilingual APE models outperform their corresponding English-Hindi and English-Marathi single-pair models by $2.5$ and $2.39$ TER points, respectively, with further notable improvements over the multilingual APE model observed through multi-task learning ($+1.29$ and $+1.44$ TER points), data augmentation ($+0.53$ and $+0.45$ TER points) and domain adaptation ($+0.35$ and $+0.45$ TER points). We release the synthetic data, code, and models accrued during this study publicly at https://github.com/cfiltnlp/Multilingual-APE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。