arXiv:2506.17080cs.CLcs.AI2025-06被引 41

让大模型同时精通翻译和通用任务,打破专精与泛用的矛盾。

Tower+: Bridging Generality and Translation Specialization in Multilingual LLMs

  • 通过多阶段训练融合预训练、监督微调与强化学习,兼顾翻译与通用能力。
  • 2B~72B模型在翻译和多语言评测中均超越同类开源及闭源模型。
  • 适合需要高精度翻译与多任务处理能力的企业级应用落地。

微调预训练大模型在机器翻译等特定任务上已证明有效,但常导致对话推理、指令遵循等通用能力下降,限制其在需多种技能融合的实际场景中的应用。本文提出Tower+系列模型,通过一种新型训练方法,在翻译专精与多语言通用能力间实现帕累托最优。该方法基于Tower(Alves et al., 2024)框架,包含持续预训练、监督微调、偏好优化及可验证奖励的强化学习。每个阶段均精心生成与筛选数据,以提升翻译、代码生成、数学解题和通用指令遵循表现。我们构建了2B、9B和72B三个规模模型。小模型性能超过更大规模的开源与闭源模型(如Llama 3.3 70B、GPT-4o)。最大模型在高资源语言翻译中达到业界领先水平,并在多语言Arena Hard评测和新提出的IF-MT基准(评估翻译与指令遵循)中取得顶尖成绩。结果表明,可在优化特定业务领域(如翻译与本地化)的同时,媲美前沿模型的通用能力。

原文摘要 · Abstract (English)

Fine-tuning pretrained LLMs has been shown to be an effective strategy for reaching state-of-the-art performance on specific tasks like machine translation. However, this process of adaptation often implies sacrificing general-purpose capabilities, such as conversational reasoning and instruction-following, hampering the utility of the system in real-world applications that require a mixture of skills. In this paper, we introduce Tower+, a suite of models designed to deliver strong performance across both translation and multilingual general-purpose text capabilities. We achieve a Pareto frontier between translation specialization and multilingual general-purpose capabilities by introducing a novel training recipe that builds on Tower (Alves et al., 2024), comprising continued pretraining, supervised fine-tuning, preference optimization, and reinforcement learning with verifiable rewards. At each stage of training, we carefully generate and curate data to strengthen performance on translation as well as general-purpose tasks involving code generation, mathematics problem solving, and general instruction-following. We develop models at multiple scales: 2B, 9B, and 72B. Our smaller models often outperform larger general-purpose open-weight and proprietary LLMs (e.g., Llama 3.3 70B, GPT-4o). Our largest model delivers best-in-class translation performance for high-resource languages and top results in multilingual Arena Hard evaluations and in IF-MT, a benchmark we introduce for evaluating both translation and instruction-following. Our findings highlight that it is possible to rival frontier models in general capabilities, while optimizing for specific business domains, such as translation and localization.

大模型多语言翻译通用能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。