arXiv:2502.09188cs.CLcs.AI2025-02NAACL被引 4

构建729亿词元的高质量波斯语语料库,推动波斯语大模型发展

Matina: A Large-Scale 73B Token Persian Text Corpus

  • 构建72.9B词元波斯语语料,经去重与清洗确保数据质量
  • 在关键NLP任务上训练的模型性能显著提升
  • 开源数据集与代码,适合波斯语NLP研究者使用

文本语料库对摘要、翻译和大语言模型(LLMs)等任务至关重要。尽管已有多种单语和多语种数据集,但波斯语因资源有限,长期处于数据匮乏状态。现有波斯语数据集普遍规模小、内容单一,主要由博客和新闻文章构成。这一高质量、多样化数据的短缺,制约了波斯语自然语言处理模型及开源大模型的发展。鉴于模型性能高度依赖训练数据质量,我们提出Matina语料库,包含72.9亿词元,经过精心预处理和去重,保障数据纯净度。我们通过在关键NLP任务上训练和评估Transformer模型,验证其有效性。该数据集及预处理代码已公开,可供研究人员持续改进,助力未来波斯语NLP进步。

原文摘要 · Abstract (English)

Text corpora are essential for training models used in tasks like summarization, translation, and large language models (LLMs). While various efforts have been made to collect monolingual and multilingual datasets in many languages, Persian has often been underrepresented due to limited resources for data collection and preprocessing. Existing Persian datasets are typically small and lack content diversity, consisting mainly of weblogs and news articles. This shortage of high-quality, varied data has slowed the development of NLP models and open-source LLMs for Persian. Since model performance depends heavily on the quality of training data, we address this gap by introducing the Matina corpus, a new Persian dataset of 72.9B tokens, carefully preprocessed and deduplicated to ensure high data quality. We further assess its effectiveness by training and evaluating transformer-based models on key NLP tasks. Both the dataset and preprocessing codes are publicly available, enabling researchers to build on and improve this resource for future Persian NLP advancements.

语料库波斯语大模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。