构建了超大规模多语言数据集,助力高精度语言模型训练
An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)
- 构建8万亿词多语言语料与3.8亿句对平行语料
- 覆盖193种语言,支持多语言模型高效训练
- 开源全流程代码,适合研究者复现与应用
训练前沿大语言模型需要海量优质文本数据。然而,构建合适的多语言数据集仍具挑战。本文提出HPLT v2,是HPLT项目扩展版的高质量单语与平行语料库集合。单语部分包含8万亿词,覆盖193种语言;平行数据包含3.8亿句对,覆盖51种语言。我们详细记录了整个数据处理流程,并开源代码以供复现。对数据质量与特性进行了广泛分析。最后,评估了基于HPLT v2训练的语言模型与机器翻译系统的性能,验证其实际价值。
原文摘要 · Abstract (English)
Training state-of-the-art large language models requires vast amounts of clean and diverse textual data. However, building suitable multilingual datasets remains a challenge. In this work, we present HPLT v2, a collection of high-quality multilingual monolingual and parallel corpora, extending prior work of the HPLT project. The monolingual portion of the data contains 8T tokens covering 193 languages, while the parallel data contains 380M sentence pairs covering 51 languages. We document the entire data pipeline and release the code to reproduce it. We provide extensive analysis of the quality and characteristics of our data. Finally, we evaluate the performance of language models and machine translation systems trained on HPLT v2, demonstrating its value.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。