arXiv:2509.14008cs.CLcs.AI2025-09被引 6

构建大规模阿拉伯语指令与翻译模型,提升本地化AI能力

Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale

  • 用压缩版双语教师模型生成高质量阿拉伯-英语对照数据
  • 训练出1.2B到90亿参数的多规模模型,在阿拉伯语任务上领先
  • 适合研究阿拉伯语NLP、多语言AI落地的开发者和学者

我们提出Hala,一套基于translate-and-tune流程构建的阿拉伯语中心指令与翻译模型。首先将强健的阿拉伯语↔英语教师模型压缩至FP8精度,实现约2倍吞吐提升且不损失质量,用于生成高保真双语监督数据。随后,使用轻量级语言模型LFM2-1.2B在此数据上微调,并将其用于将高质量英文指令集翻译为阿拉伯语,构建出百万级指令跟随专用语料库。我们训练了3.5亿、7亿、12亿和90亿参数的Hala模型,并采用slerp合并策略平衡阿拉伯语专精性与基础模型性能。在阿拉伯语基准测试中,Hala在“纳米”(≤20亿)和“小”(70亿-90亿)两类模型中均达到当前最优表现,优于其基线模型。我们已开源模型、数据、评估工具与训练方案,以推动阿拉伯语NLP研究发展。

原文摘要 · Abstract (English)

We present Hala, a family of Arabic-centric instruction and translation models built with our translate-and-tune pipeline. We first compress a strong AR$\leftrightarrow$EN teacher to FP8 (yielding $\sim$2$\times$ higher throughput with no quality loss) and use it to create high-fidelity bilingual supervision. A lightweight language model LFM2-1.2B is then fine-tuned on this data and used to translate high-quality English instruction sets into Arabic, producing a million-scale corpus tailored to instruction following. We train Hala models at 350M, 700M, 1.2B, and 9B parameters, and apply slerp merging to balance Arabic specialization with base-model strengths. On Arabic-centric benchmarks, Hala achieves state-of-the-art results within both the "nano" ($\leq$2B) and "small" (7-9B) categories, outperforming their bases. We release models, data, evaluation, and recipes to accelerate research in Arabic NLP.

阿拉伯语NLP指令模型翻译模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。