打造支持阿语和拉丁字母双书写系统的埃及方言大模型
Nile-Chat: Egyptian Language Models for Arabic and Latin Scripts
- 用分支训练混合策略融合专精脚本的专家模型
- 12B模型在拉丁字母任务上比Qwen2.5-14B高出14.4%
- 首个针对埃及方言双书写系统设计的开源大模型
我们提出Nile-Chat-4B、3x4B-A6B和12B系列大模型,专为埃及方言设计,可理解并生成阿拉伯与拉丁双书写系统文本。其中,Nile-Chat-3x4B-A6B采用创新的分支训练混合(Branch-Train-MiX)策略,将不同脚本专精的专家模型融合为单一MoE架构。在自建的埃及语评估基准上,这些模型显著优于主流多语言及阿拉伯语大模型(如LLaMa、Jais、ALLaM)。特别地,12B模型在拉丁脚本任务中相较Qwen2.5-14B-Instruct提升14.4%。所有资源均已开源。该工作为双书写系统语言的大模型适配提供了完整方法论,填补了当前大模型发展中的关键空白。
原文摘要 · Abstract (English)
We introduce Nile-Chat-4B, 3x4B-A6B, and 12B, a collection of LLMs for Egyptian dialect, uniquely designed to understand and generate texts written in both Arabic and Latin scripts. Specifically, with Nile-Chat-3x4B-A6B, we introduce a novel language adaptation approach by leveraging the Branch-Train-MiX strategy to merge script-specialized experts, into a single MoE model. Our Nile-Chat models significantly outperform leading multilingual and Arabic LLMs, such as LLaMa, Jais, and ALLaM, on our newly introduced Egyptian evaluation benchmarks, which span both understanding and generative tasks. Notably, our 12B model yields a 14.4% performance gain over Qwen2.5-14B-Instruct on Latin-script benchmarks. All our resources are publicly available. We believe this work presents a comprehensive methodology for adapting LLMs to dual-script languages, addressing an often overlooked aspect in modern LLM development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。