arXiv:2605.25420cs.CLcs.AI2026-05

四款主流大模型在索马里语安全拒绝上明显弱于英语,差距达0.40至0.93。

SomaliBench Eval: Measuring English-to-Somali Refusal Gaps in Open-Weight Language Models

论文配图:SomaliBench Eval: Measuring English-to-Somali Refusal Gaps in Open-Weight Language Models
图 1 · 摘自论文原文
  • 用双语对照的本地作者验证基准测试模型对索马里语有害请求的拒绝能力。
  • 所有模型在索马里语上的拒绝率均显著低于英语,最大差距达0.93。
  • 多数非拒绝输出是语言混乱或无关内容,而非直接危害回应,适合关注多语安全的团队。

大型语言模型的安全评估仍高度依赖英语,导致低资源语言在模型全球部署时被忽视。我们评估了四款开源指令微调模型在SomaliBench v0上的表现,该基准包含100个由本地作者验证的英-索马里语配对有害意图提示。所有模型(Llama-3.1-8B-Instruct、Gemma-2-9B-Instruct、Qwen-2.5-7B-Instruct、Aya-23-8B)在本地运行,温度设为0,使用相同的‘有益、无害、诚实’系统提示。结果显示,所有模型在英-索马里语拒绝率差距从0.40到0.93不等,且经配对自助法和精确McNemar检验均显著。其中三款模型的主要非拒绝模式并非流利危害响应,而是错误语言、逻辑混乱或偏离主题生成。一个固定版本Claude Sonnet(claude-sonnet-4-5-20250929)对每条回复分类为拒绝、遵守或模糊;其安全层共错判34次,由本地作者人工标注。本地作者抽查与裁判一致率达100%(Cohen's κ=1.00)。仅报告汇总拒绝率、类别差距与可靠性统计,原始模型输出因可能含危害内容而本地保留,未发布。

原文摘要 · Abstract (English)

Large language model safety evaluation remains heavily English-centered, leaving low-resource languages under-measured even when models are deployed globally. We evaluate four open-weight instruction-tuned models on SomaliBench v0, a native-author-verified benchmark of 100 harmful-intent prompts paired across English and Somali. Each of Llama-3.1-8B-Instruct, Gemma-2-9B-Instruct, Qwen-2.5-7B-Instruct, and Aya-23-8B is run locally with temperature 0 and the same English "helpful, harmless, and honest" (HHH) system prompt. We find large English-to-Somali refusal gaps for all four models, ranging from 0.40 to 0.93, all strictly positive under a paired bootstrap and significant by exact McNemar tests. For three models, the dominant Somali non-refusal mode is not fluent harmful compliance but unclear output: wrong-language, incoherent, or off-topic generations. A pinned Claude Sonnet snapshot (claude-sonnet-4-5-20250929) classifies each response as refused, complied, or unclear; its safety layer declined 34 of 800 classifications, which the native author labeled manually. A native-author spot-check achieves 100% agreement with the judge (Cohen's $κ=1.00$) on 74 comparable rows. We report aggregate refusal rates, category gaps, and reliability statistics only; raw model generations are retained locally and are not released because some may contain harmful content.

多语安全低资源语言模型评估索马里语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。