用轻量方法把阿拉伯语注入现有大模型,性能提升8%且不丢原有知识。
Kuwain 1.5B: An Arabic SLM via Language Injection
- 通过语言注入将阿拉伯语融入英文模型,仅用少量原数据
- 阿拉伯语多基准测试平均性能提升8%,保留原有知识能力
- 适合需低成本扩展多语言支持的场景,尤其关注阿拉伯语
提升现有模型的知识能力是人工智能发展的重要方向。本文提出一种新方法,将新语言高效融入大型语言模型(LLM)。该方法成功在不损害模型原有知识的前提下,将此前未见过的目标语言注入现有模型。我们训练了一个15亿参数的小型模型Kuwain,通过将阿拉伯语注入主要以英语训练的开源小模型实现。实验表明,该方法在多个阿拉伯语基准测试中平均性能提升8%,同时仅需极少原始模型数据即可保持其原有知识。这一方案为在英语和阿拉伯语双语环境下构建完整模型提供了低成本替代路径。结果表明,无需大规模重训或高资源投入,即可实现高效、精准的语言模型扩展。
原文摘要 · Abstract (English)
Enhancing existing models with new knowledge is a crucial aspect of AI development. This paper introduces a novel method for integrating a new language into a large language model (LLM). Our approach successfully incorporates a previously unseen target language into an existing LLM without compromising its prior knowledge. We trained a tiny model with 1.5 billion parameters named Kuwain by injecting the Arabic language into a small open-source model mainly trained in English. Our method demonstrates significant improvements in Arabic language performance, with an average 8% improvement across various benchmarks, while retaining the model's existing knowledge with a minimum amount of the original model's data. This offers a cost-effective alternative to training a comprehensive model in both English and Arabic. The results highlight the potential for efficient, targeted language model expansion without extensive retraining or resource-intensive processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。