arXiv:2411.06402cs.CLcs.AI2024-11

构建了2020亿词的阿拉伯语机器翻译语料,支持小模型训练。

Fineweb-Edu-Ar: Machine-translated Corpus to Support Arabic Small Language Models

  • 用机器翻译将英文语料转为阿拉伯语,构建大规模语料库。
  • 数据量达2020亿词,是目前公开最大的阿拉伯语机器翻译数据集。
  • 适合研究阿拉伯语小模型或低资源语言NLP的开发者使用。

随着大语言模型的发展,其对数据的需求也不断增长,尤其在多语言场景下,高质量且易获取的网络数据稀缺,促使大量合成数据集生成方法出现。其中,机器翻译(MT)是一种关键手段:将高质量英文文本转换为相对低资源的目标语言。本报告介绍FineWeb-Edu-Ar,即HuggingFace上广受欢迎的去重版FineWeb-Edu数据集的机器翻译版本。据我们所知,FineWeb-Edu-Ar是目前公开的最大规模阿拉伯语机器翻译语料库,包含2020亿词,基于阿拉伯语训练的分词器。该数据集旨在支持阿拉伯语小型语言模型的开发与研究。

原文摘要 · Abstract (English)

As large language models (LLMs) grow and develop, so do their data demands. This is especially true for multilingual LLMs, where the scarcity of high-quality and readily available data online has led to a multitude of synthetic dataset generation approaches. A key technique in this space is machine translation (MT), where high-quality English text is adapted to a target, comparatively low-resource language. This report introduces FineWeb-Edu-Ar, a machine-translated version of the exceedingly popular (deduplicated) FineWeb-Edu dataset from HuggingFace. To the best of our knowledge, FineWeb-Edu-Ar is the largest publicly available machine-translated Arabic dataset out there, with its size of 202B tokens of an Arabic-trained tokenizer.

阿拉伯语机器翻译小模型语料库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。