发现大模型生成文本中性别刻板印象持续存在,即使女性被更多提及。
More of the Same: Persistent Representational Harms Under Increased Representation
- 提出GAS(P)评估方法,检测无提示下群体表征偏差
- 实证显示不同性别描述词选择存在显著差异
- 适合关注AI公平性与偏见研究的读者
为识别和缓解生成式AI系统的危害,必须考察不同社会群体在系统中的表征方式。现有方法仅关注谁被代表,却忽略了如何被代表。本文提出GAS(P)评估方法,用于揭示无提示场景下生成文本中的群体表征偏差。将其应用于主流大语言模型的职位描述,发现尽管在生成传记时女性占比更高,但不同性别的描述仍存在显著差异:在职业语境下,男性与女性的用词分布存在统计学显著差异,且这些差异与表征伤害和刻板印象相关。结果警示,单纯增加未提示下的代表性可能反而加剧表征偏见。所提方法可系统、严谨地测量该问题。
原文摘要 · Abstract (English)
To recognize and mitigate the harms of generative AI systems, it is crucial to consider whether and how different societal groups are represented by these systems. A critical gap emerges when naively measuring or improving who is represented, as this does not consider how people are represented. In this work, we develop GAS(P), an evaluation methodology for surfacing distribution-level group representational biases in generated text, tackling the setting where groups are unprompted (i.e., groups are not specified in the input to generative systems). We apply this novel methodology to investigate gendered representations in occupations across state-of-the-art large language models. We show that, even though the gender distribution when models are prompted to generate biographies leads to a large representation of women, even representational biases persist in how different genders are represented. Our evaluation methodology reveals that there are statistically significant distribution-level differences in the word choice used to describe biographies and personas of different genders across occupations, and we show that many of these differences are associated with representational harms and stereotypes. Our empirical findings caution that naively increasing (unprompted) representation may inadvertently proliferate representational biases, and our proposed evaluation methodology enables systematic and rigorous measurement of the problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。