评测340亿参数阿拉伯语大模型在多场景下的表现,展现其语言与文化适配能力。
UI-Level Evaluation of ALLaM 34B: Measuring an Arabic-Centric LLM via HUMAIN Chat
- 通过真实对话界面测试,覆盖标准语、方言、混合语言等115次输出。
- 生成与混语任务得分高达4.92/5,方言理解达4.21/5,推理能力4.64/5。
- 适合关注阿拉伯语AI落地、多语言应用的开发者与研究者参考。
以英语为主训练的大语言模型常难以捕捉阿拉伯语的语言与文化细节。为弥补这一差距,沙特数据与人工智能局(SDAIA)推出了专注于阿拉伯语的ALLaM系列模型。其中公开可用最强大的模型是ALLaM-34B,已被HUMAIN用于构建并部署基于该模型的HUMAIN Chat闭源对话服务。本文对ALLaM-34B进行了扩展且精细化的用户界面级评估。使用涵盖现代标准阿拉伯语、五种地区方言、代码切换、事实知识、算术与时间推理、创意生成及对抗性安全性的提示集,共收集115个输出(23个提示×5轮运行),由GPT-5、Gemini 2.5 Pro、Claude Sonnet-4三名前沿大模型裁判评分。计算各类别均值(95%置信区间),分析评分分布并可视化方言维度的热力图。结果显示,生成与代码切换任务表现优异(均值4.92/5),现代标准阿拉伯语处理得分为4.74/5,推理能力达4.64/5,方言保真度提升至4.21/5,安全相关提示表现稳定可靠(4.54/5)。综合来看,这些结果表明ALLaM-34B是一个技术稳健、文化贴合的阿拉伯语大模型,具备实际部署能力。
原文摘要 · Abstract (English)
Large language models (LLMs) trained primarily on English corpora often struggle to capture the linguistic and cultural nuances of Arabic. To address this gap, the Saudi Data and AI Authority (SDAIA) introduced the $ALLaM$ family of Arabic-focused models. The most capable of these available to the public, $ALLaM-34B$, was subsequently adopted by HUMAIN, who developed and deployed HUMAIN Chat, a closed conversational web service built on this model. This paper presents an expanded and refined UI-level evaluation of $ALLaM-34B$. Using a prompt pack spanning modern standard Arabic, five regional dialects, code-switching, factual knowledge, arithmetic and temporal reasoning, creative generation, and adversarial safety, we collected 115 outputs (23 prompts times 5 runs) and scored each with three frontier LLM judges (GPT-5, Gemini 2.5 Pro, Claude Sonnet-4). We compute category-level means with 95\% confidence intervals, analyze score distributions, and visualize dialect-wise metric heat maps. The updated analysis reveals consistently high performance on generation and code-switching tasks (both averaging 4.92/5), alongside strong results in MSA handling (4.74/5), solid reasoning ability (4.64/5), and improved dialect fidelity (4.21/5). Safety-related prompts show stable, reliable performance of (4.54/5). Taken together, these results position $ALLaM-34B$ as a robust and culturally grounded Arabic LLM, demonstrating both technical strength and practical readiness for real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。