用心理刻板印象分析大模型生成故事的性别偏见,发现模型偏见可被属性条件缓解。
Investigating Gender Bias in LLM-Generated Stories via Psychological Stereotypes
- 基于心理学刻板印象设计故事生成任务,考察隐性性别偏见。
- 模型在无条件生成中男性倾向明显,但加入非性别属性可减轻偏见。
- 模型规模越大,其偏见与心理学真实分类越一致,适合伦理评估研究者参考。
随着大语言模型(LLMs)在各类应用中日益普及,其可能放大性别偏见的担忧也不断上升。以往研究多通过显式性别提示或短文本补全、问答任务探测偏见,但这些方式可能忽略长文本生成中更隐性的偏见。本文采用心理学中研究的性别刻板印象(如攻击性、爱闲聊)作为分析框架,在开放式叙事生成任务中探究模型偏见。我们构建了一个名为 StereoBias-Stories 的新数据集,包含不受限或受一个、两个或六个来自25种心理刻板印象的随机属性及三个任务相关结尾条件的故事。分析发现:(1)在无条件提示下,模型平均更偏向男性;而引入与性别刻板印象无关的属性可缓解此偏见。(2)同一性别刻板印象的多个属性叠加会加剧模型行为:男性相关属性放大偏见,女性相关属性则减轻偏见。(3)模型偏见与心理学基准分类一致,且一致性随模型规模提升。这些结果凸显了基于心理学的评估对理解模型偏见的重要性。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) are increasingly used across different applications, concerns about their potential to amplify gender biases in various tasks are rising. Prior research has often probed gender bias using explicit gender cues as counterfactual, or studied them in sentence completion and short question answering tasks. These formats might overlook more implicit forms of bias embedded in generative behavior of longer content. In this work, we investigate gender bias in LLMs using gender stereotypes studied in psychology (e.g., aggressiveness or gossiping) in an open-ended task of narrative generation. We introduce a novel dataset called StereoBias-Stories containing short stories either unconditioned or conditioned on (one, two, or six) random attributes from 25 psychological stereotypes and three task-related story endings. We analyze how the gender contribution in the overall story changes in response to these attributes and present three key findings: (1) While models, on average, are highly biased towards male in unconditioned prompts, conditioning on attributes independent from gender stereotypes mitigates this bias. (2) Combining multiple attributes associated with the same gender stereotype intensifies model behavior, with male ones amplifying bias and female ones alleviating it. (3) Model biases align with psychological ground-truth used for categorization, and alignment strength increases with model size. Together, these insights highlight the importance of psychology-grounded evaluation of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。