arXiv:2608.21985cs.AI2026-08

首个阿拉伯语大模型安全测评基准,揭示主流模型在50%攻击下失效

Redteaming Leading Arabic LLMs with ASAS

  • 构建首个全人工标注的阿拉伯语安全评测集ASAS,覆盖8类风险与8种攻击策略
  • 7款主流模型在50%危险提示前防御失败,武器与违禁品类风险最突出
  • 强调文化适配重要性,证明人类评估优于自动化工具,适合关注中东AI安全的研究者

随着大语言模型在阿拉伯语地区的广泛应用,确保其安全性和文化契合度日益重要。然而,阿拉伯语大模型的安全性,尤其是在对抗性评估场景下的研究仍不充分。本文提出首个完全由人工标注的阿拉伯语安全评测基准ASAS,包含801个涵盖8类安全风险与8种攻击策略的提示,理想回复以现代标准阿拉伯语(MSA)呈现。我们在七款具备阿拉伯语能力的领先模型(包括GPT-4o、Claude 3.7 Sonnet及地区模型ALLaM、FANAR)上进行红队测试,由人工标注者使用4分制安全量表评分。结果显示,多数模型在50%的危险提示前无法有效防御。高危害类别如武器与违禁品存在显著安全缺口,直接与隐晦攻击最为有效。结果还表明,语言对齐难以跨语言迁移,且自动化安全评估器(如GPT-4o)表现远逊于人类。ASAS为阿拉伯语大模型安全提供了一个文化贴合的评测基准与红队协议。

原文摘要 · Abstract (English)

As the adoption of large language models (LLMs) grows in Arabic-speaking regions, ensuring their safety and cultural alignment is increasingly critical. However, Arabic LLM safety remains underexplored, especially in adversarial evaluation settings. We introduce the Arabic Safety Index (ASAS), the first fully human-curated Arabic benchmark for redteaming LLMs. ASAS contains 801 prompts spanning 8 safety categories and 8 attack strategies, with ideal responses in Modern Standard Arabic (MSA). We conduct a redteaming evaluation across seven leading models with Arabic capabilities, including GPT-4o, Claude 3.7 Sonnet, and regional models such as ALLaM and FANAR. Human annotators rate responses using a structured 4-point safety scale, revealing that most models fail to defend against 50% of unsafe prompts. Our findings highlight major safety gaps in high-harm categories such as weapons and illicit substances, with direct and obfuscation-based attacks proving most effective. The results also show that language alignment does not readily transfer across languages, and that automated safety judges (e.g., GPT-4o) perform poorly compared to human annotators. ASAS provides a culturally grounded benchmark and redteaming protocol to drive progress in Arabic LLM safety.

大模型安全阿拉伯语红队测试文化对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。