arXiv:2606.21413cs.CLcs.LG2026-06

小模型专精日英互译,实测表现优于大模型。

CAT-Translate: Building Compact Open-Source Models for Japanese-English Translation

论文配图:CAT-Translate: Building Compact Open-Source Models for Japanese-English Translation
图 1 · 摘自论文原文
  • 用合成语料+两阶段微调训练0.8B~7B小模型
  • 在真实业务场景中超越多语言大模型
  • 适合资源有限但需高精度翻译的场景

当前大型多语言翻译模型在机器翻译基准测试中表现出色,这引发了一个实际问题:若仅需支持特定语对,是否值得开发专用模型?为此,我们构建了一组专用于日英双向翻译的小型语言模型(0.8B、1.4B、3.3B 和 7B 参数)。采用两阶段监督微调,并结合 Multi-Objective GRPO(Ichihara et al. 2025)方法,在合成生成的平行语料上进行训练。我们在 WMT 及涵盖商业、法律、医疗、金融和专利领域的实际应用场景中评估了这些模型。尽管多语言模型在 WMT 基准上表现优异,但我们的紧凑模型在真实世界任务中表现更优,表明即使在大模型时代,开发专用翻译模型仍具实际价值。

原文摘要 · Abstract (English)

Nowadays, large multilingual translation models demonstrate impressive translation capabilities in the machine translation benchmarks. This raises a practical question to the developers: is it worth developing translation models specialized for a particular language pair if you only need to support that language pair? To give an anecdotal answer to this question, we develop a family of small language models (0.8B, 1.4B, 3.3B, and 7B parameters) specialized for Japanese-English bidirectional translation. We employ a two-stage supervised fine-tuning approach followed by Multi-Objective GRPO (Ichihara et al. 2025) to train models on synthetically generated parallel corpora. We evaluate our models on WMT and real-world translation benchmarks across business, legal, medical, financial, and patent domains. While multilingual models achieve strong performance on WMT benchmarks, our compact models outperform them on real-world benchmarks, suggesting the practical utility of developing specialized translation models even in the era of large multilingual models.

翻译模型小模型日英翻译专用模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。