对比英斯双语模型偏见,发现偏见不转移只变形。
Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili

- 用对称英斯提示对测试模型,分析九类社会偏见。
- 偏见率最高变化12个百分点,部分模型零拒绝跨语言。
- 超55%提示生成语义不同结果,适合多语言安全研究者。
大型语言模型在多语言场景中日益普及,但安全对齐与偏见评估仍以英语为中心。我们通过向GPT-5.2和Gemini 2.5 Flash提交4,900组对称的英斯提示对,在九个社会偏见维度上生成19,600条完成内容,评估了刻板印象频率、情感倾向、拒绝行为及跨语言语义相似性。结果显示,偏见并非简单迁移而是发生转变:特定维度上的刻板印象率最高变动12个百分点,Gemini在斯瓦希里语中的中性情感率翻倍,GPT-5.2在英语中拒绝169条提示而在斯瓦希里语中拒绝0条,表明拒绝行为锚定于英语表层形式。超过55%的提示对在两个模型中生成语义差异较大的输出。这说明仅基于英语的偏见审计无法充分覆盖多语言部署风险。
原文摘要 · Abstract (English)
Large language models are increasingly deployed in multilingual contexts, yet safety alignment and bias evaluation remain overwhelmingly English-centric. We investigate whether social biases generalise across languages by submitting 4,900 symmetric English--Swahili prompt pairs to GPT-5.2 and Gemini 2.5 Flash across nine demographic bias axes, yielding 19,600 completions evaluated for stereotype prevalence, sentiment, refusal behaviour, and cross-lingual semantic similarity. Our findings show that bias transforms rather than transfers: stereotype rates shifted by up to 12 percentage points on specific axes, Gemini's neutral-sentiment rate doubled in Swahili, and GPT-5.2 refused 169 prompts in English and zero in Swahili, consistent with refusal behaviour anchored to English-language surface forms at the behavioural level. Over 55% of prompt pairs produced semantically dissimilar completions across both models. These reinforce the idea that English-only bias audits do not produce adequate coverage for multilingual deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。