arXiv:2603.04410cs.CLcs.AI2026-03

首个专为阿拉伯语大模型设计的安全评估基准,填补中文外安全评测空白。

SalamahBench: Toward Standardized Safety Evaluation for Arabic Language Models

  • 构建8170条提示的多类别安全评测集,覆盖12类危害场景。
  • 发现贾伊斯2模型在多数类别中漏洞明显,而法纳尔2表现更优但不均衡。
  • 强调需针对阿拉伯语定制安全机制,不能依赖通用英文模型判断。

语言模型的安全对齐是可信AI的基础。然而,尽管各方正推动阿拉伯语语言模型(ALMs)的应用,其系统性安全评估仍严重不足,制约了主流落地。现有安全评测基准与防护模型以英语为主,难以适用于阿拉伯语NLP系统,且掩盖了细粒度的危害类别差异。本文提出SalamahBench,一个统一的阿尔及利亚语模型安全评测基准,包含8,170个提示,覆盖MLCommons安全风险分类体系中的12个类别。通过整合异构数据集,经由人工智能过滤与多阶段人工验证的严格流程构建。使用该基准,我们评估了五种先进ALMs(包括Fanar 1/2、ALLaM 2、Falcon H1R、Jais 2),在不同防护配置下(单个防护模型、多数表决聚合、与人工标注黄金标准比对)。结果表明:尽管法纳尔2在整体攻击成功率最低,但在特定危害领域表现不一;而贾伊斯2则持续暴露高风险,说明其内在安全对齐较弱。进一步发现,原生阿拉伯语模型作为安全判别器时,性能显著劣于专用防护模型。总体表明,必须采用类别感知的评估方式和专门的防护机制,才能实现阿尔及利亚语模型的有效风险控制。

原文摘要 · Abstract (English)

Safety alignment in Language Models (LMs) is fundamental for trustworthy AI. However, while different stakeholders are trying to leverage Arabic Language Models (ALMs), systematic safety evaluation of ALMs remains largely underexplored, limiting their mainstream uptake. Existing safety benchmarks and safeguard models are predominantly English-centric, limiting their applicability to Arabic Natural Language Processing (NLP) systems and obscuring fine-grained, category-level safety vulnerabilities. This paper introduces SalamaBench, a unified benchmark for evaluating the safety of ALMs, comprising $8,170$ prompts across $12$ different categories aligned with the MLCommons Safety Hazard Taxonomy. Constructed by harmonizing heterogeneous datasets through a rigorous pipeline involving AI filtering and multi-stage human verification, SalamaBench enables standardized, category-aware safety evaluation. Using this benchmark, we evaluate five state-of-the-art ALMs, including Fanar 1 and 2, ALLaM 2, Falcon H1R, and Jais 2, under multiple safeguard configurations, including individual guard models, majority-vote aggregation, and validation against human-annotated gold labels. Our results reveal substantial variation in safety alignment: while Fanar 2 achieves the lowest aggregate attack success rates, its robustness is uneven across specific harm domains. In contrast, Jais 2 consistently exhibits elevated vulnerability, indicating weaker intrinsic safety alignment. We further demonstrate that native ALMs perform substantially worse than dedicated safeguard models when acting as safety judges. Overall, our findings highlight the necessity of category-aware evaluation and specialized safeguard mechanisms for robust harm mitigation in ALMs.

安全评测阿拉伯语大模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。