构建首个阿拉伯语大模型安全评估数据集,涵盖5799条本土化问题。
Arabic Dataset for LLM Safeguard Evaluation
- 针对阿拉伯语文化背景设计5799条安全测试题
- 发现主流模型在敏感话题上表现差异显著
- 双视角评估框架适用于政治敏感场景
大语言模型(LLMs)的安全性日益受到关注。尽管已有大量研究聚焦英语,但阿拉伯语因语言与文化复杂性,其安全性仍严重不足。本文提出一个面向阿拉伯地区、包含5,799个问题的安全评估数据集,涵盖直接攻击、间接攻击及含敏感词的无害请求,反映阿拉伯世界社会文化语境。为探究不同立场对敏感议题处理的影响,我们设计双视角评估框架,从政府与反对派两个角度评估模型响应。在五个以阿拉伯语为中心及多语言的领先模型上进行实验,结果显示其安全性能存在显著差异。这凸显了开发文化特异性数据集对于负责任部署LLMs的重要性。
原文摘要 · Abstract (English)
The growing use of large language models (LLMs) has raised concerns regarding their safety. While many studies have focused on English, the safety of LLMs in Arabic, with its linguistic and cultural complexities, remains under-explored. Here, we aim to bridge this gap. In particular, we present an Arab-region-specific safety evaluation dataset consisting of 5,799 questions, including direct attacks, indirect attacks, and harmless requests with sensitive words, adapted to reflect the socio-cultural context of the Arab world. To uncover the impact of different stances in handling sensitive and controversial topics, we propose a dual-perspective evaluation framework. It assesses the LLM responses from both governmental and opposition viewpoints. Experiments over five leading Arabic-centric and multilingual LLMs reveal substantial disparities in their safety performance. This reinforces the need for culturally specific datasets to ensure the responsible deployment of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。