用合成数据为36种欧洲语言构建了4.8万亿词的预训练语料库。
MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages

- 通过翻译1000亿高质量英文语料,生成跨语言平行数据
- 在1000亿词训练下性能超原生数据基线15%,仅需72%参数量
- 适合多语言模型研究者,尤其关注低资源语言训练
公开的大规模预训练语料仍集中于英语,制约多语言大模型发展。我们提出MultiSynt/MT,一个开源合成平行语料库,涵盖36种欧洲语言,共约4.8万亿目标语言词元,基于1000亿高质量Nemotron-CC词元,使用Tower+与OPUS-MT/HPLT-MT系统进行翻译。对众多中低资源欧洲语言而言,这是目前最大开放预训练资源。在广泛多语言基准测试中,基于MultiSynt/MT训练的参考大模型达到HPLT 2.0(原生数据基线)最终得分,仅需约72%的预训练词元;在匹配的1000亿词训练预算下,性能高出约15%。分析还揭示评估盲点:标准多项选择基准无法捕捉翻译质量差异,而以流畅性敏感的LLM作为评判者可有效识别;挪威语的习语与文化任务仍更依赖原生数据。我们发布该语料库,包含多系统生成的对齐翻译,支持多语言预训练与评估的可控研究。
原文摘要 · Abstract (English)
Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synthetic parallel corpus with approximately 4.8 trillion target-language tokens across 36 European languages, produced by translating 100 billion high-quality Nemotron-CC tokens with Tower+ and OPUS-MT/HPLT-MT systems. For many medium- and lower-resource European languages, this is the largest openly available pre-training resource. On a broad multilingual benchmark suite, reference LLMs trained on MultiSynt/MT reach the final score of HPLT 2.0, a native-data baseline, using roughly 72% fewer pre-training tokens, and outperform it by approximately 15% relative at a matched 100B-token training budget. Our analyses also identify evaluation blind spots: standard multiple-choice benchmarks miss translation-quality differences that a fluency-sensitive LLM-as-judge evaluation cleanly recovers on the trained LLMs (with no fluency deficit in MultiSynt itself), and Norwegian idiomatic and culturally grounded tasks remain better served by native data. We release the corpus, including row-aligned translations from multiple systems, to support controlled research on multilingual pre-training data and evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。