arXiv:2505.17049cs.CLcs.AI2025-05被引 10

LLM在招聘评估中普遍偏爱女性名字,且受位置影响

Gender and Positional Biases in LLM-Based Hiring Decisions: Evidence from Comparative CV/Résumé Evaluations

  • 用男女姓名对调的简历对比测试模型偏好
  • 70个职业中所有模型均更倾向女性名字候选人
  • 模型易受名字位置和性别标签影响,需谨慎用于招聘

本研究考察大型语言模型(LLMs)在评估求职者简历时的行为。实验涉及22个主流LLM,每个模型在给定一个职位描述后,被要求从一对专业匹配的简历中选择更合适的候选人,两份简历仅姓名不同(一男一女)。每对简历正反各呈现一次以消除顺序偏差。结果显示,在70种职业中,所有模型均一致倾向于女性名字候选人。当简历中添加显式性别字段(男/女)后,对女性候选人的偏好进一步增强。使用“候选人A”和“候选人B”等中性标识时,部分模型仍偏向选择“A”,而交换标识后偏好消失,实现性别平衡。单独评分时,女性简历得分略高,但差异极小。在姓名旁添加偏好代词(he/him或she/her)会略微提升人选概率,无论性别。大多数模型表现出显著的位置偏好:更倾向选择提示中排在首位的候选人。这些发现表明,部署LLM于高风险自主决策场景需谨慎,其决策未必基于原则性推理。

原文摘要 · Abstract (English)

This study examines the behavior of Large Language Models (LLMs) when evaluating professional candidates based on their resumes or curricula vitae (CVs). In an experiment involving 22 leading LLMs, each model was systematically given one job description along with a pair of profession-matched CVs, one bearing a male first name, the other a female first name, and asked to select the more suitable candidate for the job. Each CV pair was presented twice, with names swapped to ensure that any observed preferences in candidate selection stemmed from gendered names cues. Despite identical professional qualifications across genders, all LLMs consistently favored female-named candidates across 70 different professions. Adding an explicit gender field (male/female) to the CVs further increased the preference for female applicants. When gendered names were replaced with gender-neutral identifiers "Candidate A" and "Candidate B", several models displayed a preference to select "Candidate A". Counterbalancing gender assignment between these gender-neutral identifiers resulted in gender parity in candidate selection. When asked to rate CVs in isolation rather than compare pairs, LLMs assigned slightly higher average scores to female CVs overall, but the effect size was negligible. Including preferred pronouns (he/him or she/her) next to a candidate's name slightly increased the odds of the candidate being selected regardless of gender. Finally, most models exhibited a substantial positional bias to select the candidate listed first in the prompt. These findings underscore the need for caution when deploying LLMs in high-stakes autonomous decision-making contexts and raise doubts about whether LLMs consistently apply principled reasoning.

大模型招聘偏见性别偏见位置偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。