arXiv:2410.04456cs.CL2024-10被引 1

构建了超万亿词的北欧语言预训练数据集SWEb,推动小语种NLP发展。

SWEb: A Large Web Dataset for the Scandinavian Languages

  • 采用模型驱动文本提取,降低传统规则方法复杂度
  • 涵盖超1万亿词,是迄今最大北欧语言数据集
  • 专设瑞典语填空评测基准,支持模型对比验证

本文提出迄今最大的北欧语言预训练数据集——斯堪的纳维亚网络(SWEb),包含超过一万亿个标记。论文详述了数据采集与处理流程,并引入一种基于模型的文本提取方法,相比规则方法显著降低复杂度。同时,本文设计了一个新的瑞典语填空式评测基准,用于评估语言模型性能,并将SWEb训练的模型与FineWeb训练的模型进行对比,结果表现相当。所有数据、模型和代码均公开共享。

原文摘要 · Abstract (English)

This paper presents the hitherto largest pretraining dataset for the Scandinavian languages: the Scandinavian WEb (SWEb), comprising over one trillion tokens. The paper details the collection and processing pipeline, and introduces a novel model-based text extractor that significantly reduces complexity in comparison with rule-based approaches. We also introduce a new cloze-style benchmark for evaluating language models in Swedish, and use this test to compare models trained on the SWEb data to models trained on FineWeb, with competitive results. All data, models and code are shared openly.

数据集北欧语言预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。