首个支持1600种语言的机器翻译系统,突破多语言覆盖极限。
Omnilingual MT: Machine Translation for 1,600 Languages
- 融合公开语料与人工标注数据,构建超大规模多语言训练集。
- 10亿至80亿参数模型性能超越700亿参数基线,低算力下仍高效。
- 显著提升冷门语言生成质量,适合多语言应用开发者使用。
高质量机器翻译(MT)已可扩展至数百种语言,但相比全球7000种语言,当前系统仅覆盖约200种目标语言及数百种源语言,且缺乏可靠评估基准与指标。本文提出首个支持超过1600种语言的通用多语言机器翻译系统(OMT),依托整合大型公共多语言语料与新创建的人工校验数据集MeDLEY双语语料的数据策略实现。我们探索两种将大语言模型(LLM)专用于翻译的方法:作为解码器(OMT-LLaMA)或编码器-解码器架构中的模块(OMT-NLLB)。所有10亿至80亿参数模型性能均达到或优于700亿参数基线,展现明显专业化优势,可在低算力环境下保持强翻译质量。对英译1600种语言的评估显示,基线模型虽能理解冷门语言,但常无法生成有意义内容;而OMT-LLaMA模型显著扩大了可实现连贯生成的语言范围。此外,OMT模型在跨语言迁移能力上持续提升,接近解决1600种语言翻译中‘理解’环节的核心挑战。我们的排行榜及主要人工评估数据集(BOUQuET和Met-BOUQuET)正动态向全语言化演进,且免费开放。
原文摘要 · Abstract (English)
High-quality machine translation (MT) can scale to hundreds of languages, setting a high bar for multilingual systems. However, compared to the world's 7,000 languages, current systems still offer only limited coverage: about 200 languages on the target side, and maybe a few hundreds more on the source side, supported due to cross-lingual transfer. And even these numbers have been hard to evaluate due to the lack of reliable benchmarks and metrics. We present Omnilingual Machine Translation (OMT), the first MT system supporting more than 1,600 languages. This scale is enabled by a comprehensive data strategy that integrates large public multilingual corpora with newly created datasets, including manually curated MeDLEY bitext. We explore two ways of specializing a Large Language model (LLM) for machine translation: as a decoder-only model (OMT-LLaMA) or as a module in an encoder-decoder architecture (OMT-NLLB). Notably, all our 1B to 8B parameter models match or exceed the MT performance of a 70B LLM baseline, revealing a clear specialization advantage and enabling strong translation quality in low-compute settings. Moreover, our evaluation of English-to-1,600 translations further shows that while baseline models can interpret undersupported languages, they frequently fail to generate them with meaningful fidelity; OMT-LLaMA models substantially expand the set of languages for which coherent generation is feasible. Additionally, OMT models improve in cross-lingual transfer, being close to solving the "understanding" part of the puzzle in MT for the 1,600 evaluated. Our leaderboard and main human-created evaluation datasets (BOUQuET and Met-BOUQuET) are dynamically evolving towards Omnilinguality and freely available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。