研究发现大模型对亚裔和西裔的刻板印象更严重,且在不同设置下依然存在。
How Robust Is Homogeneity Bias in LLMs? Evidence Across Models, Decoding Settings, and Identity Signals
- 测试7个大模型在多种推理设置下的群体同质化偏见
- 亚裔和西裔被描述得比白人更同质,这一现象稳定存在
- 用姓氏代替标签会反转黑人群体的偏见方向,说明表达方式影响结果
大型语言模型存在同质化偏见——即倾向于将边缘群体描绘得比主流群体内部更相似。本研究在七个开源指令微调模型(参数量7-20B)、5×5的温度与top-p组合解码网格,以及两种群体身份提示方式(显式标签与具有种族特征的姓氏)下系统评估该偏见。在七种模型中的六种中,亚裔与西裔美国人默认配置下被描述为显著比白人更同质,且在所有温度与top-p条件下平均仍呈正向;非裔与性别偏见则在不同模型间方向不一。保守的单元级再分析确认了亚裔与西裔同质化偏见的稳健性,而多数非裔与性别信号未通过检验。在基于姓氏的身份提示范式中,亚裔与西裔同质化偏见再次显现,但带有黑人特征的姓氏引发的输出反而比白人姓氏更少同质——这种反向效应在标签范式中不存在,表明身份表达方式深刻影响哪些偏见浮现及其方向。
原文摘要 · Abstract (English)
Large language models (LLMs) reproduce homogeneity bias -- the tendency to portray marginalized groups as more internally similar than dominant groups -- but whether this bias generalizes across models, is stable under different inference settings, or depends on how group identity is signaled remains unstudied. We map homogeneity bias across seven open-weight instruction-tuned LLMs (7-20B parameters), a 5x5 temperature x top-p decoding grid, and two paradigms for signaling group identity (explicit labels vs. racially distinctive names). In six of seven models, Hispanic and Asian Americans are portrayed as significantly more homogeneous than White Americans at the default configuration, and the effect remains positive on average at every temperature and top-p tested; African American and gender bias instead vary in direction across models. A conservative cell-level re-analysis confirms Hispanic and Asian homogeneity as robust while weaker African American and gender signals largely do not survive, establishing group-specific robustness. We also apply the same grid to a names-based paradigm in which group identity is signaled via racially distinctive surnames rather than explicit labels. The names paradigm corroborates Hispanic and Asian homogeneity bias, but Black-coded surnames elicit robustly less homogeneous outputs than White-coded names in every model tested -- a reversal absent from the label paradigm -- showing that how group identity is operationalized shapes which biases surface and in which direction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。