Juhaina是专为阿拉伯语使用者设计的双语大模型,更懂阿拉伯文化。
CamelEval: Advancing Culturally Aligned Arabic Language Models and Benchmarks
- 基于92.4亿参数和8192上下文窗口,专注阿拉伯语文化对齐
- 在阿拉伯语回答、区域事实准确性和文化理解上优于同类模型
- 开源发布,适合需要本地化AI的中东研究与应用
大型语言模型(LLMs)是现代人工智能系统的核心。本文介绍Juhaina,一个专为阿拉伯语使用者价值观和偏好设计的阿拉伯语-英语双语大模型,具备指令遵循、开放问答、信息提供和文本处理等高级功能。该模型包含92.4亿参数,训练上下文窗口最大达8192个标记。本文详述了Juhaina的构建过程,并进行了广泛的实证评估。此外,我们指出现有广泛采用的Open Arabic LLM Leaderboard(OALL)存在局限性,提出新的评估基准CamelEval。结果表明,Juhaina在生成阿拉伯语有用回应、提供区域事实准确性信息以及理解细微文化特征方面,超越了同等规模的Llama和Gemma系列模型。我们希望Juhaina能为超过4亿阿拉伯语使用者普及前沿AI技术,提供不仅会说其语言,更能理解其文化的模型。所有模型已公开发布于Huggingface(https://huggingface.co/elmrc)。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are the cornerstones of modern artificial intelligence systems. This paper introduces Juhaina, a Arabic-English bilingual LLM specifically designed to align with the values and preferences of Arabic speakers. Juhaina inherently supports advanced functionalities such as instruction following, open-ended question answering, information provisioning, and text processing. Our model contains 9.24 billion parameters and is trained on a context window of up to 8,192 tokens. This paper details the creation process of Juhaina and provides an extensive empirical evaluation. Furthermore, we identify the limitations of widely-adopted Open Arabic LLM Leaderboard (OALL) and propose a new evaluation benchmark, CamelEval. Our findings demonstrate that Juhaina surpasses existing LLMs of comparable sizes, such as the Llama and Gemma families, in generating helpful responses in Arabic, providing factually accurate information about the region, and understanding nuanced cultural aspects. We aspire for Juhaina to democratize cutting-edge AI technologies, serving over 400 million Arabic speakers by offering LLMs that not only communicate in their language but also comprehend their culture. We publicly release all models on Huggingface \url{https://huggingface.co/elmrc}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。