arXiv:2511.16204cs.LGstat.ME2025-11

用因果生成模型合成招聘数据,提升算法公平性。

Causal Synthetic Data Generation in Recruitment

  • 构建两个因果生成模型,分别模拟职位和简历数据
  • 生成的合成数据能控制偏差,验证排序公平性
  • 适合研究招聘算法公平性的研究人员使用

合成数据生成(SDG)在数据质量差或受隐私与监管限制的领域愈发重要。招聘领域因简历信息(如性别、残疾状况、年龄)敏感,公开数据稀缺,制约了公平透明机器学习模型的发展,尤其是依赖大量数据的候选人排序算法。当前因果生成模型(CGM)提供新路径:可生成保留原始因果关系的合成数据,增强数据生成过程的公平性与可解释性。本文提出一种专用SDG方法,包含两个CGM:一个建模职位信息,一个建模简历内容。两者基于领域知识构建因果图,生成合成数据集,并在引入特定偏差的可控场景中评估候选者排序的公平性。

原文摘要 · Abstract (English)

The importance of Synthetic Data Generation (SDG) has increased significantly in domains where data quality is poor or access is limited due to privacy and regulatory constraints. One such domain is recruitment, where publicly available datasets are scarce due to the sensitive nature of information typically found in curricula vitae, such as gender, disability status, or age. This lack of accessible, representative data presents a significant obstacle to the development of fair and transparent machine learning models, particularly ranking algorithms that require large volumes of data to effectively learn how to recommend candidates. In the absence of such data, these models are prone to poor generalisation and may fail to perform reliably in real-world scenarios. Recent advances in Causal Generative Models (CGMs) offer a promising solution. CGMs enable the generation of synthetic datasets that preserve the underlying causal relationships within the data, providing greater control over fairness and interpretability in the data generation process. In this study, we present a specialised SDG method involving two CGMs: one modelling job offers and the other modelling curricula. Each model is structured according to a causal graph informed by domain expertise. We use these models to generate synthetic datasets and evaluate the fairness of candidate rankings under controlled scenarios that introduce specific biases.

合成数据因果模型招聘算法公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。