arXiv:2604.06421cs.CL2026-04

开源阿拉伯语大模型突破性能瓶颈,靠稀疏专家路由与文化适配推理蒸馏。

State-of-the-Art Arabic Language Modeling with Sparse MoE Fine-Tuning and Chain-of-Thought Distillation

  • 用稀疏MoE架构+分阶段思维链蒸馏,融合阿拉伯语言特征与地区伦理规范。
  • 在7项基准上均达SOTA,Grammar任务超越GPT-5.1显著表现,安全与检索任务也领先。
  • 适合关注低资源语言、文化敏感模型与低成本高效适配的研究者与开发者。

本文提出阿拉伯语开源大模型 Arabic-DeepSeek-R1,采用稀疏MoE骨干网络,应对欠代表语言的数字公平问题,并在全部Open Arabic LLM Leaderboard(OALL)基准上取得新SOTA。其四阶段思维链蒸馏方案整合了阿拉伯语特定语言验证与区域伦理规范,基于372M tokens、80/20阿英混合训练数据(污染可控)。该模型在七项基准测试中平均得分最高,尤其在语法导向的MadinahQA任务上显著超越GPT-5.1与当前榜首;在安全导向的AraTrust、多能力任务AlGhafa及增强检索的ALRAGE上亦达SOTA或近SOTA。结果表明,稀疏MoE架构结合文化驱动的思维链蒸馏与战略性双语数据筛选,使开源适配模型系统性超越专有前沿系统GPT-5.1,在多数评估语言特异性任务中实现首次突破。这说明当前阿拉伯语大模型性能差距主要源于专业度不足而非架构限制,参数高效适配开放推理模型可无须工业级预训练即达成顶尖表现。Arabic-DeepSeek-R1建立了可复现的主权与领域特定语言技术框架,证明针对低资源语言,战略化、文化根基的稀疏MoE适配是达成标杆性能的可行且经济路径。

原文摘要 · Abstract (English)

This paper introduces Arabic-DeepSeek-R1, an application-driven open-source Arabic LLM that leverages a sparse MoE backbone to address the digital equity gap for under-represented languages, and establishes a new SOTA across the entire Open Arabic LLM Leaderboard (OALL). Our four-phase CoT distillation scheme integrates Arabic-specific linguistic verification and regional ethical norms into a 372M-token, contamination-controlled 80/20 Arabic-English training mixture. Arabic-DeepSeek-R1 achieves the highest average score across the seven-benchmark OALL suite while establishing SOTA or near-SOTA, including dominant results on grammar-focused MadinahQA (surpassing both GPT-5.1 and the OALL leader by substantial margins), safety-oriented AraTrust, multi-ability AlGhafa, and retrieval-augmented ALRAGE. Our results indicate that the combination of sparse MoE architecture, culturally-informed CoT distillation with explicit Arabic linguistic checks, and strategic bilingual data curation enables an open-source adapted model to systematically outperform the proprietary frontier system GPT-5.1 on the majority of benchmarks evaluating comprehensive language-specific tasks: the first such demonstration for Arabic LLMs. These findings indicate that much of Arabic's performance deficit in current LLM ecosystems stems from under-specialization rather than architectural limitations, and that parameter-efficient adaptation of open reasoning models can yield breakthrough SOTA performance without industrial-scale pretraining costs. Arabic-DeepSeek-R1 establishes a validated and replicable framework for sovereign and domain-specific language technologies, demonstrating that strategic, culturally-grounded adaptation of sparse MoE backbones offers a viable and cost-effective pathway to achieving record-breaking performance across standardized benchmarks for low-resource languages.

阿拉伯语稀疏MoE思维链蒸馏低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。