arXiv:2505.14256cs.CLcs.AI2025-05被引 1

FuxiMT通过稀疏化大模型,提升中文主导的多语言翻译性能。

FuxiMT: Sparsifying Large Language Models for Chinese-Centric Multilingual Machine Translation

  • 采用两阶段训练:先在大规模中文语料预训练,再在65种语言的平行数据上微调。
  • 在低资源场景下显著超越主流模型,零样本翻译能力突出。
  • 适合需要跨语言沟通、缺乏平行语料的场景,如小语种翻译。

本文提出FuxiMT,一种基于稀疏化大语言模型(LLM)的中文主导多语言机器翻译模型。采用两阶段策略:首先在大规模中文语料上预训练,随后在涵盖65种语言的大型平行数据集上进行多语言微调。FuxiMT融合专家混合(Mixture-of-Experts, MoEs)机制,并采用课程学习策略以提升在不同资源水平下的鲁棒性。实验结果表明,FuxiMT显著优于多个强基线模型,包括顶尖的LLM和机器翻译系统,尤其在低资源场景下表现优异。此外,该模型对未见语言对展现出卓越的零样本翻译能力,显示出在缺乏平行数据时弥合沟通鸿沟的巨大潜力。

原文摘要 · Abstract (English)

In this paper, we present FuxiMT, a novel Chinese-centric multilingual machine translation model powered by a sparsified large language model (LLM). We adopt a two-stage strategy to train FuxiMT. We first pre-train the model on a massive Chinese corpus and then conduct multilingual fine-tuning on a large parallel dataset encompassing 65 languages. FuxiMT incorporates Mixture-of-Experts (MoEs) and employs a curriculum learning strategy for robust performance across various resource levels. Experimental results demonstrate that FuxiMT significantly outperforms strong baselines, including state-of-the-art LLMs and machine translation models, particularly under low-resource scenarios. Furthermore, FuxiMT exhibits remarkable zero-shot translation capabilities for unseen language pairs, indicating its potential to bridge communication gaps where parallel data are scarce or unavailable.

多语言翻译稀疏化中文中心零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。