研究发现大模型招聘系统对不同人群存在显著偏见,尤其在简历检索阶段。
Small Changes, Large Consequences: Analyzing the Allocational Fairness of LLMs in Hiring Contexts
- 通过控制变量的合成数据测试模型对不同性别和种族的响应差异。
- 种族变化比性别变化更易引发摘要内容差异,且排名对两类变化均敏感。
- 检索阶段模型对非人口属性变化也敏感,提示公平性问题或源于模型脆弱性。
大型语言模型(LLMs)在招聘等高风险场景中应用日益广泛,但其生成与检索设置下的不公平决策仍缺乏研究。本文通过简历摘要生成和申请人排序两个真实人力资源任务,分析了基于LLM的招聘系统的分配公平性。我们构建了带有受控扰动的合成简历数据集,并整理了职位发布信息,考察模型在不同人口群体间的响应差异。结果显示,生成摘要在种族扰动下出现显著差异的频率高于性别扰动;模型在不同人口群体间表现出非均匀的检索选择模式,并对性别与种族扰动均呈现高度排名敏感性。令人意外的是,检索模型对非人口属性变化也表现出类似敏感度,表明公平性问题可能源自模型整体脆弱性。总体而言,基于LLM的招聘系统,特别是在检索阶段,可能产生明显偏见并导致现实中的歧视性结果。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly being deployed in high-stakes applications like hiring, yet their potential for unfair decision-making remains understudied in generative and retrieval settings. In this work, we examine the allocational fairness of LLM-based hiring systems through two tasks that reflect actual HR usage: resume summarization and applicant ranking. By constructing a synthetic resume dataset with controlled perturbations and curating job postings, we investigate whether model behavior differs across demographic groups. Our findings reveal that generated summaries exhibit meaningful differences more frequently for race than for gender perturbations. Models also display non-uniform retrieval selection patterns across demographic groups and exhibit high ranking sensitivity to both gender and race perturbations. Surprisingly, retrieval models can show comparable sensitivity to both demographic and non-demographic changes, suggesting that fairness issues may stem from broader model brittleness. Overall, our results indicate that LLM-based hiring systems, especially in the retrieval stage, can exhibit notable biases that lead to discriminatory outcomes in real-world contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。