无需假设数据分布,用统计方法精准识别模型偏见
Empirical Likelihood-Based Fairness Auditing: Distribution-Free Certification and Flagging
- 基于经验似然构建非参数化偏差检测机制
- 在COMPAS数据上发现非裔男性青年误判率显著偏高
- 计算效率提升数个数量级,适合大规模模型审计
高风险场景下的机器学习模型常在敏感子群体间表现出系统性性能差异,引发算法偏见关注。公平性审计通过认证(验证是否满足公平约束)和标记(识别遭受不公平对待的群体)来应对这一风险。然而现有方法常受限于严格的分布假设或高昂计算开销。本文提出一种基于经验似然(EL)的新框架,构建鲁棒的模型性能差异统计量。该方法为非参数化,其偏差统计量渐近服从卡方或混合卡方分布,无需假设底层数据分布即可实现有效推断。框架采用约束优化轮廓,具有稳定的数值解,支持大规模认证与高效子群体发现。实验表明,相比基于自助法的方法,EL方法覆盖率更接近名义水平,计算延迟降低数个数量级。在COMPAS数据集上,成功识别出交叉偏见:25岁以下非裔男性正向预测率显著高于群体均值,而白人女性则存在系统性低估。
原文摘要 · Abstract (English)
Machine learning models in high-stakes applications, such as recidivism prediction and automated personnel selection, often exhibit systematic performance disparities across sensitive subpopulations, raising critical concerns regarding algorithmic bias. Fairness auditing addresses these risks through two primary functions: certification, which verifies adherence to fairness constraints; and flagging, which isolates specific demographic groups experiencing disparate treatment. However, existing auditing techniques are frequently limited by restrictive distributional assumptions or prohibitive computational overhead. We propose a novel empirical likelihood-based (EL) framework that constructs robust statistical measures for model performance disparities. Unlike traditional methods, our approach is non-parametric; the proposed disparity statistics follow asymptotically chi-square or mixed chi-square distributions, ensuring valid inference without assuming underlying data distributions. This framework uses a constrained optimization profile that admits stable numerical solutions, facilitating both large-scale certification and efficient subpopulation discovery. Empirically, the EL methods outperform bootstrap-based approaches, yielding coverage rates closer to nominal levels while reducing computational latency by several orders of magnitude. We demonstrate the practical utility of this framework on the COMPAS dataset, where it successfully flags intersectional biases, specifically identifying a significantly higher positive prediction rate for African-American males under 25 and a systemic under-prediction for Caucasian females relative to the population mean.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。