arXiv:2511.01066cs.CL2025-11被引 5

构建30万亿词多语言数据集,支持190+语言的LLM训练与评估

HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models

  • 从网络爬取构建30万亿词多语言数据,含语言识别、去重、质量标注等全流程开源管道
  • 覆盖9个欧洲语言的评测基准,支持原生任务与抗提示敏感性评估,模型性能显著提升
  • 提供57个单语编码器-解码器模型及大量平行语料,适合多语言NLP研究者使用

我们启动一项持续计划,为近200种语言提供开放、超大规模、高质量且丰富标注的文本数据集。总规模达30万亿词,可能是目前可用的最大通用多语言LLM预训练数据集合。数据源自多个来源的网络爬取,配套完整的开源流水线:从网页存档中选择文档、从HTML提取文本、对噪声文本进行语言识别、精确与近似去重、标注包括语域标签、文本质量估计和个人身份信息;最终完成筛选与过滤。通过对比与分析统计、对24种语言样本的人工抽检,以及基于该数据训练的多种语言模型架构的端到端评估,验证了数据质量。在多语言大模型评估方面,我们提供涵盖九种欧洲语言的综合性基准,特别关注原生任务设计、缓解提示敏感性的机制,以及评分的精细化归一化与聚合。此外,我们训练并评估了一组57个单语编码器-解码器模型,以及若干单语GPT类参考模型。除单语数据与模型外,还展示了从该数据中自动挖掘的大规模平行文本,并提出一种通过机器翻译合成的新颖平行语料。

原文摘要 · Abstract (English)

We present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely the largest generally available multilingual collection of LLM pre-training data. These datasets are derived from web crawls from different sources and accompanied with a complete, open-source pipeline for document selection from web archives, text extraction from HTML, language identification for noisy texts, exact and near-deduplication, annotation with, among others, register labels, text quality estimates, and personally identifiable information; and final selection and filtering. We report on data quality probes through contrastive and analytical statistics, through manual inspection of samples for 24 languages, and through end-to-end evaluation of various language model architectures trained on this data. For multilingual LLM evaluation, we provide a comprehensive collection of benchmarks for nine European languages, with special emphasis on natively created tasks, mechanisms to mitigate prompt sensitivity, and refined normalization and aggregation of scores. Additionally, we train and evaluate a family of 57 monolingual encoder-decoder models, as well as a handful of monolingual GPT-like reference models. Besides the monolingual data and models, we also present a very large collection of parallel texts automatically mined from this data, together with a novel parallel corpus synthesized via machine translation.

多语言数据集大模型预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。