arXiv:2409.17912cs.CL2024-09被引 28

首个专为摩洛哥方言设计的LLM,显著提升本地语言理解与生成能力。

Atlas-Chat: Adapting Large Language Models for Low-Resource Moroccan Arabic Dialect

  • 整合真实数据与合成数据构建指令微调集,针对摩洛哥方言优化训练
  • 9B模型在方言评估集上比更大13B模型高13%性能,超越主流阿拉伯语LLM
  • 开源全部资源,为低资源语言提供可复用的指令微调方法

我们提出Atlas-Chat,首个专为方言阿拉伯语设计的大语言模型系列。聚焦摩洛哥阿拉伯语(即Darija),通过整合现有资源、人工与合成方式创建新数据集,并经严格质量控制翻译英文指令,构建指令微调数据集。基于该数据集微调的Atlas-Chat-2B、9B和27B模型,在遵循方言指令及标准NLP任务中表现优异。尤其在我们新推出的涵盖判别与生成任务的DarijaMMLU评估套件中,9B模型相较更大13B模型提升13%性能,优于现有顶尖及阿拉伯语专用LLM(如LLaMa、Jais、AceGPT)。此外,我们分析不同微调策略与基线模型选择,确定最优配置。所有资源公开可用,本工作为低资源语言的指令微调提供了完整方法论,填补当前大模型普遍忽视低资源语言的空白。

原文摘要 · Abstract (English)

We introduce Atlas-Chat, the first-ever collection of LLMs specifically developed for dialectal Arabic. Focusing on Moroccan Arabic, also known as Darija, we construct our instruction dataset by consolidating existing Darija language resources, creating novel datasets both manually and synthetically, and translating English instructions with stringent quality control. Atlas-Chat-2B, 9B, and 27B models, fine-tuned on the dataset, exhibit superior ability in following Darija instructions and performing standard NLP tasks. Notably, our models outperform both state-of-the-art and Arabic-specialized LLMs like LLaMa, Jais, and AceGPT, e.g., our 9B model gains a 13% performance boost over a larger 13B model on DarijaMMLU, in our newly introduced evaluation suite for Darija covering both discriminative and generative tasks. Furthermore, we perform an experimental analysis of various fine-tuning strategies and base model choices to determine optimal configurations. All our resources are publicly accessible, and we believe our work offers comprehensive design methodologies of instruction-tuning for low-resource languages, which are often neglected in favor of data-rich languages by contemporary LLMs.

方言识别大模型低资源语言指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。