arXiv:2409.14842cs.AIcs.CL2024-09被引 3

华为用多种训练策略提升机器翻译性能,大模型后处理显著改进多领域译文质量。

HW-TSC's Submission to the CCMT 2024 Machine Translation Tasks

  • 采用正则化丢弃、双向训练等10种策略优化Transformer-big模型
  • 多领域任务中结合大模型后编辑,译文质量显著提升
  • 适合关注工业级翻译系统优化的研究者与开发者

本文介绍华为翻译服务研究中心(HW-TSC)在第二十届中国机器翻译大会(CCMT 2024)中的机器翻译任务参赛方案。参与双语机器翻译和多领域机器翻译两项任务。针对这两项任务,基于深度Transformer-big架构,采用正则化丢弃、双向训练、数据多样化、前向翻译、反向翻译、交替训练、课程学习及自训练集成学习等多种训练策略训练神经机器翻译(NMT)模型。为进一步探索大语言模型(LLM)对NMT系统翻译质量的提升作用,我们使用监督微调方法,将Llama2-13b训练为自动后编辑(APE)模型,用于优化多领域任务中的NMT输出结果。通过这些综合策略,我们的提交在最终评估中取得具有竞争力的成绩。

原文摘要 · Abstract (English)

This paper presents the submission of Huawei Translation Services Center (HW-TSC) to machine translation tasks of the 20th China Conference on Machine Translation (CCMT 2024). We participate in the bilingual machine translation task and multi-domain machine translation task. For these two translation tasks, we use training strategies such as regularized dropout, bidirectional training, data diversification, forward translation, back translation, alternated training, curriculum learning, and transductive ensemble learning to train neural machine translation (NMT) models based on the deep Transformer-big architecture. Furthermore, to explore whether large language model (LLM) can help improve the translation quality of NMT systems, we use supervised fine-tuning to train llama2-13b as an Automatic post-editing (APE) model to improve the translation results of the NMT model on the multi-domain machine translation task. By using these plyometric strategies, our submission achieves a competitive result in the final evaluation.

机器翻译大模型应用神经网络后编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。