arXiv:2507.02152cs.AI2025-07

用审计研究数据发现公平干预的虚假公平,揭示10%隐藏歧视

The Illusion of Fairness: Auditing Fairness Interventions with Audit Studies

  • 用虚构测试者做随机对照实验获取高质量歧视数据
  • 传统公平干预看似公平,实则仍存在约10%的隐性偏差
  • 基于个体处理效应的新方法能进一步降低算法歧视

人工智能系统,尤其是机器学习模型,正被广泛应用于招聘、贷款发放等复杂决策领域。评估这些AI系统及其人类决策对应物的有效性与公平性,是计算和社会科学共同关注的重要课题。在机器学习中,常通过重采样训练数据来缓解下游分类器中的偏差,例如使不同受保护群体的招聘率在训练集中均等化。然而,这类方法通常仅在便利样本数据上评估,引入了选择偏差和标签偏差。在社会科学中,审计研究通过随机对照试验向目标(如职位空缺、商家、医生)发送虚构的‘测试者’(如简历、邮件、患者演员),可提供高质量数据以精确估计歧视程度。本文探讨如何利用审计研究数据改进自动化招聘算法的训练与评估。结果表明,此类数据揭示了常见公平干预方法在传统指标下看似实现平等,但经适当测量后仍存在约10%的差异。我们还提出了基于个体处理效应估计的新干预方法,进一步降低了算法歧视。

原文摘要 · Abstract (English)

Artificial intelligence systems, especially those using machine learning, are being deployed in domains from hiring to loan issuance in order to automate these complex decisions. Judging both the effectiveness and fairness of these AI systems, and their human decision making counterpart, is a complex and important topic studied across both computational and social sciences. Within machine learning, a common way to address bias in downstream classifiers is to resample the training data to offset disparities. For example, if hiring rates vary by some protected class, then one may equalize the rate within the training set to alleviate bias in the resulting classifier. While simple and seemingly effective, these methods have typically only been evaluated using data obtained through convenience samples, introducing selection bias and label bias into metrics. Within the social sciences, psychology, public health, and medicine, audit studies, in which fictitious ``testers'' (e.g., resumes, emails, patient actors) are sent to subjects (e.g., job openings, businesses, doctors) in randomized control trials, provide high quality data that support rigorous estimates of discrimination. In this paper, we investigate how data from audit studies can be used to improve our ability to both train and evaluate automated hiring algorithms. We find that such data reveals cases where the common fairness intervention method of equalizing base rates across classes appears to achieve parity using traditional measures, but in fact has roughly 10% disparity when measured appropriately. We additionally introduce interventions based on individual treatment effect estimation methods that further reduce algorithmic discrimination using this data.

公平性审计算法歧视审计研究机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。