开源多语言翻译模型,专攻中文与少数民族语言互译。
Hunyuan-MT Technical Report
- 采用分阶段训练,融合预训练、微调与强化学习提升性能。
- 在31个语言对中30个第一,尤其擅长中文与少数民族语言翻译。
- 提出类慢思考模型,融合多参数输出,效果优于传统方法。
本文介绍 Hunyuan-MT-7B,首个开源的多语言翻译模型,支持33种主要语言间的双向翻译,特别关注中文与若干少数民族语言及方言之间的互译。为应对多样化翻译场景并提升测试时表现,我们提出受慢思考模式启发的 Hunyuan-MT-Chimera-7B 模型,该模型整合 Hunyuan-MT-7B 在不同参数设置下生成的多个输出,实现超越传统基于思维链(CoT)的慢思考模型的表现。模型开发遵循针对多语言翻译优化的整体训练流程,包含通用预训练、以机器翻译为导向的预训练、监督微调(SFT),以及通过强化学习(RL)和弱到强强化学习实现的高级对齐。全面实验表明,Hunyuan-MT-7B 和 Hunyuan-MT-Chimera-7B 在同等参数量的专用翻译模型中显著领先,并超越多数当前最优大模型,尤其在中文与少数民族语言及方言互译任务中表现突出。在 WMT2025 共享任务(通用机器翻译)中,模型在31个语言对中的30个排名第一,验证了其在高资源语言(如中文、英文、日文)和低资源语言(如捷克语、马拉地语、爱沙尼亚语、冰岛语)上的强大泛化能力。
原文摘要 · Abstract (English)
In this report, we introduce Hunyuan-MT-7B, our first open-source multilingual translation model, which supports bidirectional translation across 33 major languages and places a special emphasis on translation between Mandarin and several ethnic minority languages as well as dialects. Furthermore, to serve and address diverse translation scenarios and enhance model performance at test time, we introduce Hunyuan-MT-Chimera-7B, a translation model inspired by the slow thinking mode. This model integrates multiple outputs generated by the Hunyuan-MT-7B model under varying parameter settings, thereby achieving performance superior to that of conventional slow-thinking models based on Chain-of-Thought (CoT). The development of our models follows a holistic training process specifically engineered for multilingual translation, which begins with general and MT-oriented pre-training to build foundational capabilities, proceeds to Supervised Fine-Tuning (SFT) for task-specific adaptation, and culminates in advanced alignment through Reinforcement Learning (RL) and weak-to-strong RL. Through comprehensive experimentation, we demonstrate that both Hunyuan-MT-7B and Hunyuan-MT-Chimera-7B significantly outperform all translation-specific models of comparable parameter size and most of the SOTA large models, particularly on the task of translation between Mandarin and minority languages as well as dialects. In the WMT2025 shared task (General Machine Translation), our models demonstrate state-of-the-art performance, ranking first in 30 out of 31 language pairs. This result highlights the robustness of our models across a diverse linguistic spectrum, encompassing high-resource languages such as Chinese, English, and Japanese, as well as low-resource languages including Czech, Marathi, Estonian, and Icelandic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。