arXiv:2409.15924cs.CLcs.AI2024-09被引 2

华为用多语言迁移等策略提升西班牙小语种翻译效果。

Multilingual Transfer and Domain Adaptation for Low-Resource Languages of Spain

  • 采用多语言迁移与回译增强小语种模型训练。
  • 在西-阿拉贡语等任务中取得有竞争力的翻译性能。
  • 适合关注低资源语言翻译的技术团队参考。

本文介绍了华为翻译服务中心(HW-TSC)在WMT 2024年西班牙低资源语言翻译任务中的参赛情况。我们参与了三种翻译任务:西班牙语到阿拉贡语(es-arg)、西班牙语到阿兰语(es-arn)以及西班牙语到阿斯图里亚语(es-ast)。针对这些任务,我们基于深度Transformer-big架构的神经机器翻译(NMT)模型,采用了多语言迁移、正则化丢弃、前向翻译与反向翻译、LaBSE去噪、转导集成学习等训练策略。通过上述优化方法,我们的提交结果在最终评估中表现优异,达到具有竞争力的水平。

原文摘要 · Abstract (English)

This article introduces the submission status of the Translation into Low-Resource Languages of Spain task at (WMT 2024) by Huawei Translation Service Center (HW-TSC). We participated in three translation tasks: spanish to aragonese (es-arg), spanish to aranese (es-arn), and spanish to asturian (es-ast). For these three translation tasks, we use training strategies such as multilingual transfer, regularized dropout, forward translation and back translation, labse denoising, transduction ensemble learning and other strategies to neural machine translation (NMT) model based on training deep transformer-big architecture. By using these enhancement strategies, our submission achieved a competitive result in the final evaluation.

机器翻译低资源语言多语言迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。