arXiv:2410.23825cs.CLcs.AI2024-10NeurIPS被引 20

构建首个覆盖千种语言的开源清洁语料库,支持小语种研究。

GlotCC: An Open Broad-Coverage CommonCrawl Corpus and Pipeline for Minority Languages

  • 基于CommonCrawl数据,通过可复现的开源流水线生成
  • 1000+语言覆盖,2TB文档级语料,严格去噪
  • 适合小语种NLP、语言资源建设者使用

预训练语言模型的发展推动了对大规模文本语料的需求,尤其在模型缩放定律发现后更为迫切。现有语料库大多仅覆盖主流语言,缺乏同时满足三大条件的资源:(i)覆盖广泛的小语种;(ii)基于开源可复现的生成流程;(iii)经过严格去噪,具备可信性。本文提出GlotCC,一个从CommonCrawl提取的2TB通用领域文档级语料库,涵盖超过1000种语言。我们公开发布GlotCC v.1.0及生成系统——包括管道、语言识别模型和过滤器——版本3.0可在GitHub获取。语料库地址:https://huggingface.co/datasets/cis-lmu/GlotCC-v1,管道代码:https://github.com/cisnlp/GlotCC。

原文摘要 · Abstract (English)

The need for large text corpora has increased with the advent of pretrained language models and, in particular, the discovery of scaling laws for these models. Most available corpora have sufficient data only for languages with large dominant communities. However, there is no corpus available that (i) covers a wide range of minority languages; (ii) is generated by an open-source reproducible pipeline; and (iii) is rigorously cleaned from noise, making it trustworthy to use. We present GlotCC, a clean, document-level, 2TB general domain corpus derived from CommonCrawl, covering more than 1000 languages. We make GlotCC and the system used to generate it - including the pipeline, language identification model, and filters - available to the research community. Corpus v. 1.0 https://huggingface.co/datasets/cis-lmu/GlotCC-v1, Pipeline v. 3.0 https://github.com/cisnlp/GlotCC.

语料库小语种开源语言识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。