arXiv:2507.11548cs.CYcs.AI2025-07被引 2

AI简历筛选工具看似公平,实则可能根本不会评估能力。

Fairness Is Not Enough: Auditing Competence and Intersectional Bias in AI-powered Resume Screening

  • 用虚构简历测试八款AI工具,发现其评估能力不足
  • 部分模型对特定群体有隐性惩罚,且表现不一致
  • 建议同时审计公平性与实际评估能力,避免虚假公正

将生成式AI用于简历筛选常被视作比人类判断更少偏见,但这一前提未解决关键问题:这些系统是否具备有效评估的能力。本研究对八款主流简历筛选AI平台进行了双重审计。基于‘中立错觉’概念,实验1通过匹配虚构简历测试种族与性别偏见,发现偏见以情境依赖和交叉形式存在;某些模型因识别出人口统计信号而惩罚候选人,另一些则在不同角色与身份下表现不一致。实验2通过简历-职位不匹配及关键词操纵测试评估任务胜任力,发现部分表面无偏的模型无法区分与岗位相关或无关的经验。此时输出看似中立,实为系统未能完成有意义评估。因此,仅评估公平性不足。论文提出双验证框架,要求同时审计偏见与评估能力,作为负责任部署的最低标准。

原文摘要 · Abstract (English)

The use of publicly available generative AI systems for resume evaluation is often justified by the assumption that these tools reduce bias relative to human judgment. However, this framing leaves a prior question unresolved: whether these systems are capable of performing the evaluative task at all. This study presents a two-part audit of eight widely used AI platforms used for resume screening. Drawing on the concept of the Illusion of Neutrality, the study examines cases in which systems appear demographically unbiased because they lack the ability to meaningfully differentiate among candidates. Experiment 1 evaluates racial and gender bias using matched fictitious resumes and finds that bias persists in context-dependent and intersectional forms. Some models penalize candidates for the presence of demographic signals, while others exhibit inconsistent patterns across roles and identities under controlled conditions. Experiment 2 evaluates task competence using resume-role mismatch and keyword-manipulation tests and finds that several models that appear relatively unbiased fail to distinguish relevant from irrelevant candidate experience in relation to the target role. In these cases, outputs appear stable or neutral not because the systems are fair, but because they fail to perform meaningful evaluation. Fairness assessments alone are therefore insufficient. The paper proposes a dual-validation framework that requires auditing for both demographic bias and evaluative competence as a minimum condition for responsible deployment.

AI偏见简历筛选评估能力公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。