测试大模型在西非低资源语言中的安全拒绝能力,发现中文拒绝率下降超一半。
LSR: Linguistic Safety Robustness Benchmark for Low-Resource West African Languages
- 用中英文对照探针对比测试模型安全拒绝行为
- 西非语言拒绝率仅35%-55%,伊加拉语下降最严重
- 提出新指标量化跨语言安全退化,适合安全研究者使用
大语言模型的安全对齐主要依赖英语数据。当有害意图用低资源语言表达时,英语中有效的拒绝机制常失效。我们提出LSR(语言安全鲁棒性)基准,首个系统评估西非语言(约鲁巴语、豪萨语、伊博语、伊加拉语)中跨语言拒绝退化问题。采用双探针评估协议,将匹配的英语与目标语言探针输入同一模型,并引入拒绝中心漂移(RCD)指标,量化模型在目标语言中丢失的英语拒绝行为。我们在14个文化相关攻击探针上评估Gemini 2.5 Flash,英语拒绝率约90%,西非语言中降至35%-55%,其中伊加拉语退化最严重(RCD=0.55)。LSR已集成至Inspect AI评估框架,可公开贡献至UK AISI的inspect_evals仓库,参考实现与数据集均开放可用。
原文摘要 · Abstract (English)
Safety alignment in large language models relies predominantly on English-language training data. When harmful intent is expressed in low-resource languages, refusal mechanisms that hold in English frequently fail to activate. We introduce LSR (Linguistic Safety Robustness), the first systematic benchmark for measuring cross-lingual refusal degradation in West African languages: Yoruba, Hausa, Igbo, and Igala. LSR uses a dual-probe evaluation protocol - submitting matched English and target-language probes to the same model - and introduces Refusal Centroid Drift (RCD), a metric that quantifies how much of a model's English refusal behavior is lost when harmful intent is encoded in a target language. We evaluate Gemini 2.5 Flash across 14 culturally grounded attack probes in four harm categories. English refusal rates hold at approximately 90 percent. Across West African languages, refusal rates fall to 35-55 percent, with Igala showing the most severe degradation (RCD = 0.55). LSR is implemented in the Inspect AI evaluation framework and is available as a PR-ready contribution to the UK AISI's inspect_evals repository. A live reference implementation and the benchmark dataset are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。