arXiv:2510.13481cs.LG2025-10

构建阿拉伯语大模型的数据与训练指南,解决语言特殊性难题

Tahakom LLM Guidelines and Recipes: From Pre-training Data to an Arabic LLM

  • 系统化筛选阿拉伯语预训练数据,提升数据质量
  • 对比不同分词器对模型性能的影响,优化文本表示
  • 改进现有评估框架,适合研究者复现与协作

大型语言模型显著推动了自然语言处理的发展,提升了多领域语言理解与生成能力。然而,阿拉伯语大模型的构建面临独特挑战。本文聚焦数据筛选、分词器设计与评估体系三大关键环节,详述阿拉伯语预训练数据的采集与过滤方法,分析不同分词器对模型表现的影响,并指出现有阿拉伯语评估框架的局限性,提出系统的修正方案。为促进透明与协作,本文公开数据与方法,助力阿拉伯语语言建模发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have significantly advanced the field of natural language processing, enhancing capabilities in both language understanding and generation across diverse domains. However, developing LLMs for Arabic presents unique challenges. This paper explores these challenges by focusing on critical aspects such as data curation, tokenizer design, and evaluation. We detail our approach to the collection and filtration of Arabic pre-training datasets, assess the impact of various tokenizer designs on model performance, and examine the limitations of existing Arabic evaluation frameworks, for which we propose a systematic corrective methodology. To promote transparency and facilitate collaborative development, we share our data and methodologies, contributing to the advancement of language modeling, particularly for the Arabic language.

大模型阿拉伯语数据筛选分词器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。