用故事生成发现:模型女性角色多,但更符合性别刻板印象。
More Women, Same Stereotypes: Unpacking the Gender Bias Paradox in Large Language Models
- 用自由叙事测试模型性别偏见,揭示深层认知偏差。
- 十款主流模型均高估女性在职业中的比例,但更贴近刻板印象。
- 适合关注AI公平性、伦理与社会影响的研究者和从业者。
大型语言模型(LLMs)虽推动了自然语言处理发展,但其仍可能反映或放大社会偏见。本研究提出一种新评估框架,通过自由形式的讲故事任务揭示模型中的性别偏见。对十款主流LLMs的系统分析显示,女性在职业角色中被过度代表,这可能源于监督微调(SFT)和基于人类反馈的强化学习(RLHF)。然而,讽刺的是,尽管女性占比偏高,这些模型生成的职业性别分布反而比真实劳动力数据更接近人类刻板印象。这凸显了制定平衡缓解措施的重要性,以促进公平并防止产生新的偏见。相关提示和生成故事已发布于GitHub。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have revolutionized natural language processing, yet concerns persist regarding their tendency to reflect or amplify social biases. This study introduces a novel evaluation framework to uncover gender biases in LLMs: using free-form storytelling to surface biases embedded within the models. A systematic analysis of ten prominent LLMs shows a consistent pattern of overrepresenting female characters across occupations, likely due to supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF). Paradoxically, despite this overrepresentation, the occupational gender distributions produced by these LLMs align more closely with human stereotypes than with real-world labor data. This highlights the challenge and importance of implementing balanced mitigation measures to promote fairness and prevent the establishment of potentially new biases. We release the prompts and LLM-generated stories at GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。