arXiv:2507.02087cs.LGcs.CL2025-07被引 13

专用招聘模型比通用大模型更准更公平,能兼顾招聘效果与公平性。

Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions

  • 用领域定制模型匹配候选人,优于通用大模型。
  • 定制模型准确率AUC达0.85,种族影响比最低为0.957(接近平等)。
  • 实证表明高准确率与公平性可兼得,适合高风险决策场景。

大型语言模型(LLMs)在招聘中有望简化筛选流程,但若缺乏充分保障,可能引发准确性与算法偏见问题。本文评估了OpenAI、Anthropic、Google、Meta和Deepseek等多家机构的前沿基础模型,并与自研的领域专用招聘模型Match Score进行对比,针对约1万条真实候选人-职位配对数据,衡量其预测准确率(ROC AUC、PR AUC、F1)与公平性(基于性别、种族及交叉群体的截断影响比)。结果表明,Match Score在准确率上超越通用模型(ROC AUC 0.85 vs 0.77),并在不同群体间实现更均衡的结果:种族影响比最低达0.957(近平等),而最优通用模型仅为0.809;交叉群体影响比分别为0.906与0.773。研究指出,预训练偏见可能导致通用模型在招聘中延续社会偏见,而经过监督训练的定制模型能更有效缓解此类问题。结论强调,在高风险领域如招聘中,必须采用领域专用建模与偏见审计,警惕直接使用现成大模型的风险。同时实证显示,精准与公平并非对立,合理设计的算法可兼顾二者。

原文摘要 · Abstract (English)

The use of large language models (LLMs) in hiring promises to streamline candidate screening, but it also raises serious concerns regarding accuracy and algorithmic bias where sufficient safeguards are not in place. In this work, we benchmark several state-of-the-art foundational LLMs - including models from OpenAI, Anthropic, Google, Meta, and Deepseek, and compare them with our proprietary domain-specific hiring model (Match Score) for job candidate matching. We evaluate each model's predictive accuracy (ROC AUC, Precision-Recall AUC, F1-score) and fairness (impact ratio of cut-off analysis across declared gender, race, and intersectional subgroups). Our experiments on a dataset of roughly 10,000 real-world recent candidate-job pairs show that Match Score outperforms the general-purpose LLMs on accuracy (ROC AUC 0.85 vs 0.77) and achieves significantly more equitable outcomes across demographic groups. Notably, Match Score attains a minimum race-wise impact ratio of 0.957 (near-parity), versus 0.809 or lower for the best LLMs, (0.906 vs 0.773 for the intersectionals, respectively). We discuss why pretraining biases may cause LLMs with insufficient safeguards to propagate societal biases in hiring scenarios, whereas a bespoke supervised model can more effectively mitigate these biases. Our findings highlight the importance of domain-specific modeling and bias auditing when deploying AI in high-stakes domains such as hiring, and caution against relying on off-the-shelf LLMs for such tasks without extensive fairness safeguards. Furthermore, we show with empirical evidence that there shouldn't be a dichotomy between choosing accuracy and fairness in hiring: a well-designed algorithm can achieve both accuracy in hiring and fairness in outcomes.

招聘智能大模型应用公平性评估偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。