arXiv:2505.02609cs.AIcs.CY2025-05

研究偏见数据对招聘算法选优能力的影响,揭示算法如何受歧视与自我审查干扰。

Study of the influence of a biased database on the prediction of standard algorithms for selecting the best candidate for an interview

  • 构建模拟外部歧视与内部自审的偏见数据集用于训练
  • 五种经典算法在偏见数据下表现下降,最优候选识别率显著降低
  • 匿名化文件能缓解偏见影响,但无法完全消除算法偏差

人工智能被广泛应用于招聘流程中,以自动筛选最佳候选人,企业宣称其过程无偏见。然而,这些算法或由人工标注训练,或基于存在偏见的历史经验学习。本文通过生成模拟外部歧视和内部自审(自我审查)的偏见数据,训练五种经典算法,并评估其在客观标准下识别最优候选人的能力。同时,研究了简历匿名化对预测质量的影响。结果表明,偏见数据显著降低了算法的选优准确率,而匿名化虽有一定改善作用,但无法从根本上消除算法偏差。

原文摘要 · Abstract (English)

Artificial intelligence is used at various stages of the recruitment process to automatically select the best candidate for a position, with companies guaranteeing unbiased recruitment. However, the algorithms used are either trained by humans or are based on learning from past experiences that were biased. In this article, we propose to generate data mimicking external (discrimination) and internal biases (self-censorship) in order to train five classic algorithms and to study the extent to which they do or do not find the best candidates according to objective criteria. In addition, we study the influence of the anonymisation of files on the quality of predictions.

招聘算法偏见数据公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。