7B小模型突破阿拉伯语企业应用瓶颈,兼顾文化敏感与指令遵循。
Command R7B Arabic: A Small, Enterprise Focused, Multilingual, and Culturally Aware Arabic LLM
- 用合成数据+人工校验扩充阿拉伯语训练集
- 迭代微调后在多项阿拉伯语评测中超越同类模型
- 适合需要文化感知的中东企业级应用开发
由于数字化阿拉伯语数据有限,构建适用于企业场景的高质量大语言模型仍具挑战。本文提出一种数据合成与精炼策略,通过合成数据生成和人工标注扩展阿拉伯语训练语料。同时设计了关键的迭代后训练流程,有效对齐人类偏好,满足企业级应用需求。最终推出一个70亿参数、开源权重的小型模型,其在头对头对比及涵盖文化知识、指令遵循、RAG和上下文忠实性的阿拉伯语专项评测中表现优于同类模型。
原文摘要 · Abstract (English)
Building high-quality large language models (LLMs) for enterprise Arabic applications remains challenging due to the limited availability of digitized Arabic data. In this work, we present a data synthesis and refinement strategy to help address this problem, namely, by leveraging synthetic data generation and human-in-the-loop annotation to expand our Arabic training corpus. We further present our iterative post training recipe that is essential to achieving state-of-the-art performance in aligning the model with human preferences, a critical aspect to enterprise use cases. The culmination of this effort is the release of a small, 7B, open-weight model that outperforms similarly sized peers in head-to-head comparisons and on Arabic-focused benchmarks covering cultural knowledge, instruction following, RAG, and contextual faithfulness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。