用反事实生成技术检测AI面试评估中的偏见,揭示不同人群评分差异。
Behind the Screens: Uncovering Bias in AI-Driven Video Interview Assessments Using Counterfactuals
- 通过生成对抗网络改变申请人属性,构造反事实样本进行公平性分析。
- 发现不同性别、族裔群体在人格预测中存在显著评分差异。
- 适用于黑箱AI招聘平台的可扩展审计工具,提升算法透明度。
AI增强的人格评估正日益影响招聘决策,利用情感计算从大五人格(OCEAN)模型预测特质。然而,将AI引入此类评估引发伦理担忧,尤其训练数据中的偏见可能放大歧视,导致基于性别、族裔、年龄等受保护属性的不公平结果。为此,我们提出一种基于反事实的框架,系统评估并量化AI人格评估中的偏见。该方法采用生成对抗网络(GANs)生成改变受保护属性的申请人反事实表示,实现无需访问底层模型的公平性分析。与传统仅关注单模态或静态数据的偏见评估不同,本方法支持视觉、音频和文本多模态评估。在先进人格预测模型上的应用揭示了不同人口群体间的显著差异。我们还通过受保护属性分类器验证反事实生成的有效性。该工作为商业AI招聘平台的公平性审计提供可扩展工具,尤其适用于训练数据和模型内部不可访问的黑箱场景。结果强调反事实方法在提升情感计算伦理透明度中的重要性。
原文摘要 · Abstract (English)
AI-enhanced personality assessments are increasingly shaping hiring decisions, using affective computing to predict traits from the Big Five (OCEAN) model. However, integrating AI into these assessments raises ethical concerns, especially around bias amplification rooted in training data. These biases can lead to discriminatory outcomes based on protected attributes like gender, ethnicity, and age. To address this, we introduce a counterfactual-based framework to systematically evaluate and quantify bias in AI-driven personality assessments. Our approach employs generative adversarial networks (GANs) to generate counterfactual representations of job applicants by altering protected attributes, enabling fairness analysis without access to the underlying model. Unlike traditional bias assessments that focus on unimodal or static data, our method supports multimodal evaluation-spanning visual, audio, and textual features. This comprehensive approach is particularly important in high-stakes applications like hiring, where third-party vendors often provide AI systems as black boxes. Applied to a state-of-the-art personality prediction model, our method reveals significant disparities across demographic groups. We also validate our framework using a protected attribute classifier to confirm the effectiveness of our counterfactual generation. This work provides a scalable tool for fairness auditing of commercial AI hiring platforms, especially in black-box settings where training data and model internals are inaccessible. Our results highlight the importance of counterfactual approaches in improving ethical transparency in affective computing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。