用多语言辩论数据集暴露大模型隐性偏见,发现英语外语言偏见更严重
Surfacing Subtle Stereotypes: A Multilingual, Debate-Oriented Evaluation of Modern LLMs
- 设计多语言辩论任务,模拟真实对话中叙事偏见的生成
- 77%非洲相关表述被刻板化为落后,阿拉伯人与恐怖主义关联超89%
- 低资源语言中偏见加剧,说明现有对齐方法无法全球通用
大型语言模型广泛用于开放对话,但现有偏见评估多限于英文分类任务。本文提出 exttt{corpusname},一个涵盖四个敏感领域(女性权利、落后性、恐怖主义、宗教)的多语言辩论式评测基准,覆盖七种语言(从高资源英语、中文到低资源斯瓦希里语、尼日利亚皮钦语),共8,400个结构化辩论提示。使用GPT-4o、Claude 3.5 Haiku、DeepSeek-Chat和LLaMA-3-70B四款主流模型生成超10万条回应,并自动识别哪些群体被赋予刻板角色或现代角色。结果显示,所有模型虽经安全对齐仍再现深层偏见:阿拉伯人与恐怖主义/宗教关联≥89%,非洲人被刻板化为经济落后(最高达77%),西方群体始终被塑造为现代进步。偏见在低资源语言中显著加剧,表明以英语为主训练的对齐策略无法跨语言泛化。研究揭示当前对齐方法虽降低显性毒性,却无法避免开放场景中的系统性偏见。我们公开 exttt{corpusname} 基准与分析框架,推动下一代多语言公平评估与文化包容性对齐。
原文摘要 · Abstract (English)
Large language models (LLMs) are widely deployed for open-ended communication, yet most bias evaluations still rely on English, classification-style tasks. We introduce \corpusname, a new multilingual, debate-style benchmark designed to reveal how narrative bias appears in realistic generative settings. Our dataset includes 8{,}400 structured debate prompts spanning four sensitive domains -- Women's Rights, Backwardness, Terrorism, and Religion -- across seven languages ranging from high-resource (English, Chinese) to low-resource (Swahili, Nigerian Pidgin). Using four flagship models (GPT-4o, Claude~3.5~Haiku, DeepSeek-Chat, and LLaMA-3-70B), we generate over 100{,}000 debate responses and automatically classify which demographic groups are assigned stereotyped versus modern roles. Results show that all models reproduce entrenched stereotypes despite safety alignment: Arabs are overwhelmingly linked to Terrorism and Religion ($\geq$89\%), Africans to socioeconomic ``backwardness'' (up to 77\%), and Western groups are consistently framed as modern or progressive. Biases grow sharply in lower-resource languages, revealing that alignment trained primarily in English does not generalize globally. Our findings highlight a persistent divide in multilingual fairness: current alignment methods reduce explicit toxicity but fail to prevent biased outputs in open-ended contexts. We release our \corpusname benchmark and analysis framework to support the next generation of multilingual bias evaluation and safer, culturally inclusive model alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。