用合成简历构建公平招聘算法测试数据集
Synthetic CVs To Build and Test Fairness-Aware Hiring Tools
- 基于真实捐赠数据生成1730份带背景特征的合成简历
- 提供可衡量偏见的基准数据集,支持公平性评估
- 适合研究算法歧视与公平性验证的学者使用
算法招聘在某些领域日益重要,能高效处理数百甚至上千份简历。然而,这类系统常因简历(CV)表示不当而引入偏见,导致年龄、性别或国籍等非相关因素影响筛选结果。为开发和评估公平性技术,需包含多样化背景特征的简历数据集,但此类数据目前缺失。本文提出一种方法,利用数据捐赠收集的真实材料生成合成简历数据集,并公开1730份简历,旨在成为算法招聘中歧视研究的基准参考。
原文摘要 · Abstract (English)
Algorithmic hiring has become increasingly necessary in some sectors as it promises to deal with hundreds or even thousands of applicants. At the heart of these systems are algorithms designed to retrieve and rank candidate profiles, which are usually represented by Curricula Vitae (CVs). Research has shown, however, that such technologies can inadvertently introduce bias, leading to discrimination based on factors such as candidates' age, gender, or national origin. Developing methods to measure, mitigate, and explain bias in algorithmic hiring, as well as to evaluate and compare fairness techniques before deployment, requires sets of CVs that reflect the characteristics of people from diverse backgrounds. However, datasets of these characteristics that can be used to conduct this research do not exist. To address this limitation, this paper introduces an approach for building a synthetic dataset of CVs with features modeled on real materials collected through a data donation campaign. Additionally, the resulting dataset of 1,730 CVs is presented, which we envision as a potential benchmarking standard for research on algorithmic hiring discrimination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。