检测大模型面试回答中的性别偏见,发现其与职业性别刻板印象高度一致。
Gender Bias in LLM-generated Interview Responses
- 对比GPT-3.5、GPT-4、Claude在多种职位和问题类型下的生成结果。
- 发现男性角色回答更强调领导力,女性角色更倾向表达合作性,符合传统性别刻板印象。
- 适用于关注AI公平性的开发者、招聘平台及伦理审查者。
大语言模型(LLMs)已成为辅助生成各类文本(包括求职相关文本)的有力工具,但其生成内容中的性别偏见日益凸显。本研究对GPT-3.5、GPT-4和Claude三款模型进行了多维度审计,涵盖不同模型、问题类型和职位,评估其生成面试回答与两种性别刻板印象的一致性。结果表明,性别偏见具有持续性,且与性别刻板印象及职位社会主导性密切相关。该研究为系统性审视大模型生成面试回答中的性别偏见提供了依据,强调在相关应用中需采取审慎策略以缓解此类偏见。
原文摘要 · Abstract (English)
LLMs have emerged as a promising tool for assisting individuals in diverse text-generation tasks, including job-related texts. However, LLM-generated answers have been increasingly found to exhibit gender bias. This study evaluates three LLMs (GPT-3.5, GPT-4, Claude) to conduct a multifaceted audit of LLM-generated interview responses across models, question types, and jobs, and their alignment with two gender stereotypes. Our findings reveal that gender bias is consistent, and closely aligned with gender stereotypes and the dominance of jobs. Overall, this study contributes to the systematic examination of gender bias in LLM-generated interview responses, highlighting the need for a mindful approach to mitigate such biases in related applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。