1.6B参数的阿拉伯语大模型,小体积却性能强劲。
Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic
- 基于Stable LM 2 1.6B微调,专为阿拉伯语优化。
- 在多个基准上表现超越参数量8倍大的模型。
- 合成指令数据提升微调效果,适合资源受限场景。
大型语言模型(LLM)在自然语言处理多个领域表现出色,但主要集中在英语。近年来,更多模型加入了多语言文本以支持低资源语言。在阿拉伯语NLP领域,过去两年出现了多个专注于阿拉伯语的模型,并在多个基准上取得显著成果。然而,大多数阿拉伯语大模型参数超过70亿,导致硬件需求高、推理延迟大。本文提出阿拉伯语版Stable LM 1.6B,包含基础版和对话版,是小型但强大的阿拉伯语专用模型。其对话版本在多个基准测试中表现优异,超越了参数量高达8倍的其他模型。此外,通过引入大规模合成对话数据进行指令微调,进一步提升了性能。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown impressive results in multiple domains of natural language processing (NLP) but are mainly focused on the English language. Recently, more LLMs have incorporated a larger proportion of multilingual text to represent low-resource languages. In Arabic NLP, several Arabic-centric LLMs have shown remarkable results on multiple benchmarks in the past two years. However, most Arabic LLMs have more than 7 billion parameters, which increases their hardware requirements and inference latency, when compared to smaller LLMs. This paper introduces Arabic Stable LM 1.6B in a base and chat version as a small but powerful Arabic-centric LLM. Our Arabic Stable LM 1.6B chat model achieves impressive results on several benchmarks beating multiple models with up to 8x the parameters. In addition, we show the benefit of mixing in synthetic instruction tuning data by augmenting our fine-tuning data with a large synthetic dialogue dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。