arXiv:2605.30804cs.CL2026-05

跨语言审计发现大模型性别偏见比人类群体差异大2.5倍

Anchoring LLM Gender Bias to Human Baselines: A Cross-Lingual Audit

论文配图:Anchoring LLM Gender Bias to Human Baselines: A Cross-Lingual Audit
图 1 · 摘自论文原文
  • 用人类跨文化数据作基准,量化模型性别刻板印象偏离程度
  • 英语主导模型在韩语提示下偏见强度达本地基准的5倍
  • 揭示翻译会重构刻板印象属性,非简单缩放

我们对六种大型语言模型(LLMs)在英语、韩语、中文和日语中的性别刻板印象进行了跨语言审计。其中三种主要面向英语用户(Claude、GPT、Gemini),三种面向东亚用户(DeepSeek、Syn-Pro、HyperCLOVA X)。采用HEXACO-100人格量表,并以覆盖48个国家的跨文化人类数据集为基准,不问模型是否偏见,而是考察其性别归因与部署人群之间的偏离程度。结果显示,模型的刻板印象跨度约为人类跨国家差异的2.5倍,且跨语言效应可叠加。一个英语主导模型在韩语提示下,偏见程度达到本地基准的5倍,即使提示中已说明候选人已被录用——这本应削弱人类刻板印象。为描述此类行为而不进行排名,我们提出四类模式框架:一致性、抑制、重组、放大,覆盖24个(模型×语言)组合。项目级分析表明,翻译不仅缩放刻板印象,还改变其关联属性,表面看似校准,实则存在显著重构。研究最终表明,单一去偏方案难以在语言边界间均衡有效。

原文摘要 · Abstract (English)

We audit six large language models (LLMs) for gender stereotyping across English, Korean, Chinese, and Japanese. Three were developed primarily for English-language use (Claude, GPT, Gemini) and three for East Asian use (DeepSeek, Syn-Pro, HyperCLOVA X). We adopt the HEXACO-100 personality inventory and anchor each model against a cross-cultural human dataset spanning 48 countries to ask not whether LLMs are biased, but how far their gender attributions drift from the populations they are deployed among. Our findings show that their stereotyping spans a range roughly 2.5 times wider than the entire cross-country range found in humans, and the effect can compound across languages. One English-centric model, prompted in Korean, reached 5 times the local baseline, even when the prompt stated the candidate had already been hired, which often dampens human stereotyping. To characterize such behaviors without ranking them, we introduce a four-pattern framework -- concordance, suppression, reorganization, and amplification -- across 24 (model x language) cells. Item-level analysis reveals that translation does not just rescale stereotypes, but changes the attributes tied to it, hiding significant rearrangement under the surface while appearing well-calibrated. Our results ultimately suggest that no single debiasing pipeline is likely to address bias evenly across linguistic boundaries.

大模型偏见跨语言性别刻板去偏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。