14个大模型招聘筛选测试发现,近年模型从偏白转为偏黑,性别也类似。
Can LLMs Hire Fairly? Racial Bias in Resume Screening
- 用配对简历法测试14个主流大模型的招聘偏见
- 2023年模型仍有+2.12%白人优势,2024年后全转为黑人优势(最高-3.01%)
- 趋势表明算法招聘偏见随模型迭代发生方向性逆转
我们采用Kline、Rose和Walters(2022)的配对简历方法,审计了14个主流大语言模型在招聘中的歧视问题。唯一2023年发布的模型重现了实地实验中记录的白人回调优势(+2.12个百分点,1%显著性水平)。所有2024年及以后发布的模型均显示零差距或显著的黑人优势(最高达-3.01个百分点)。该模式在性别维度上同样成立。基于每模型24,024对职位数据,共14个模型的分析揭示了算法招聘偏见随模型代际更迭出现方向性逆转。
原文摘要 · Abstract (English)
We audit fourteen mainstream large language models (LLMs) for hiring discrimination using the paired-resume methodology of Kline, Rose, and Walters (2022). The sole 2023-vintage model reproduces the pro-White callback gap documented in field experiments on labor market discrimination ($+2.12$ pp, significant at the 1\% level). Every model released in 2024 or after shows either a null gap or a significant pro-Black reversal (up to $-3.01$ pp). The same pattern holds on the gender axis. Based on 24,024 paired postings per model across 14 models, our results document a reversal in the direction of algorithmic hiring bias across model generations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。