发现大模型在非英语语种中安全对齐失效,可能加剧文化偏见。
Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs

- 构建多语言评测基准INCLUDE,覆盖6种印度语言
- 开放源代码模型在孟加拉语中偏见最高,英文则相反
- 揭示闭源模型在英语中反而更易产生偏见,适合本地化部署研究者
当前大型语言模型的安全对齐训练严重依赖英语。当这些安全过滤器在非英语语言中失效时,语音助手和对话系统可能生成强化刻板印象的内容,绕过以英语为中心的安全机制,向非英语使用者传播有害偏见。对于印度多语言人口而言,这构成关键缺陷。为此,我们提出INCLUDE(印度文化视角的偏见理解与检测),一个用于量化印度本土社会文化偏见的多语言评估基准。INCLUDE包含2,604个提示,涵盖英语、印地语、孟加拉语、马拉地语、泰米尔语和印地英混语(Hinglish)六种语言。我们评估了十款开源与闭源大模型,分析14,988个偏见评分。统计结果显示:第一,开放源代码模型在孟加拉语中平均偏见得分最高;第二,英语在开放源代码模型中偏见最低,但在闭源模型中反而偏见最高。
原文摘要 · Abstract (English)
Current safety alignment training for Large Language Models (LLMs) are heavily English-centric. When such safety filters fail for non-English languages, the consequences are immediate and user-facing: voice assistants and spoken dialogue systems may produce stereotype-reinforcing outputs, bypassing the standard English-focused safety alignments and propagating harmful bias to non-English speaking communities. For spoken language technologies deployed across India's linguistically diverse population, this represents a critical failure mode. To address this cross-lingual gap, we introduce INCLUDE (Indian Cultural Lens for Understanding and Detecting Embedded Biases), a multilingual evaluation benchmark designed to quantify Indian-centric socio-cultural biases. INCLUDE consists of 2,604 prompts spanning six prompt languages: English, Hindi, Bengali, Marathi, Tamil, and Hinglish (Hindi-English code-mix). We evaluate ten open- and closed-source LLMs against this benchmark, analyzing 14,988 bias scores. Our statistical results reveal two key findings. First, Bengali yielded the highest average bias score in open-source models. Second, English demonstrated a notable reversal, producing the lowest bias in open-source models but the highest bias in closed-source models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。