构建简历评估公平性基准,检测大模型在招聘中的种族性别偏见
FAIRE: Assessing Racial and Gender Bias in AI-Driven Resume Evaluations
- 设计双方法测试:直接打分与排序,对比不同身份伪装简历表现
- 发现所有模型均有偏见,程度与方向差异显著
- 开源数据集与代码,适合研究者评估招聘AI公平性
随着AI驱动的招聘日益普及,公平性与偏见问题愈发关键。本文提出FAIRE(简历评估公平性评估)基准,用于检验大语言模型在跨行业简历评估中是否存在种族与性别偏见。采用直接评分和排序两种方法,通过微调简历以体现不同种族或性别身份,考察模型性能变化。结果表明,尽管所有模型均存在一定程度的偏见,但其强度与方向差异显著。该基准为系统评估模型公平性提供了清晰路径,并凸显了降低招聘类AI偏见的紧迫性。相关代码与数据集已开源:https://github.com/athenawen/FAIRE-Fairness-Assessment-In-Resume-Evaluation.git。
原文摘要 · Abstract (English)
In an era where AI-driven hiring is transforming recruitment practices, concerns about fairness and bias have become increasingly important. To explore these issues, we introduce a benchmark, FAIRE (Fairness Assessment In Resume Evaluation), to test for racial and gender bias in large language models (LLMs) used to evaluate resumes across different industries. We use two methods-direct scoring and ranking-to measure how model performance changes when resumes are slightly altered to reflect different racial or gender identities. Our findings reveal that while every model exhibits some degree of bias, the magnitude and direction vary considerably. This benchmark provides a clear way to examine these differences and offers valuable insights into the fairness of AI-based hiring tools. It highlights the urgent need for strategies to reduce bias in AI-driven recruitment. Our benchmark code and dataset are open-sourced at our repository: https://github.com/athenawen/FAIRE-Fairness-Assessment-In-Resume-Evaluation.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。