针对小群体公平性检测难题,提出自适应检验框架,让判断更可靠。
Size-adaptive Hypothesis Testing for Fairness
- 大群体用中心极限定理构造统计检验,保证假阳性率可控。
- 小群体用贝叶斯狄利克雷-多项式模型,任意样本量下都可得可信区间。
- 适合做多维敏感属性交叉分析的公平性评估,尤其数据稀疏时有用。
判断算法是否对特定人口群组存在歧视,通常通过比较公平性指标点估计与预设阈值实现。但此做法统计脆弱:忽略抽样误差,且对大小不同的子群体一视同仁。在交叉分析中,多个敏感属性联合考虑会生成更多更小的群体,数据愈发稀疏,导致公平性指标置信区间过宽,难以得出有效结论。本文提出统一的、规模自适应的假设检验框架,将公平性评估转为基于证据的统计决策。贡献有二:(i) 对足够大的子群体,证明了统计均等差异的中心极限定理,可得解析置信区间和类型一错误控制在α水平的Wald检验;(ii) 对长尾的小交叉群体,采用完全贝叶斯狄利克雷-多项式估计器,蒙特卡洛可信区间可适配任意样本量,并随数据增加自然收敛至Wald区间。我们在基准数据集上验证该方法,展示其在不同数据可用性和交叉程度下仍能提供可解释、统计严谨的决策。
原文摘要 · Abstract (English)
Determining whether an algorithmic decision-making system discriminates against a specific demographic typically involves comparing a single point estimate of a fairness metric against a predefined threshold. This practice is statistically brittle: it ignores sampling error and treats small demographic subgroups the same as large ones. The problem intensifies in intersectional analyses, where multiple sensitive attributes are considered jointly, giving rise to a larger number of smaller groups. As these groups become more granular, the data representing them becomes too sparse for reliable estimation, and fairness metrics yield excessively wide confidence intervals, precluding meaningful conclusions about potential unfair treatments. In this paper, we introduce a unified, size-adaptive, hypothesis-testing framework that turns fairness assessment into an evidence-based statistical decision. Our contribution is twofold. (i) For sufficiently large subgroups, we prove a Central-Limit result for the statistical parity difference, leading to analytic confidence intervals and a Wald test whose type-I (false positive) error is guaranteed at level $α$. (ii) For the long tail of small intersectional groups, we derive a fully Bayesian Dirichlet-multinomial estimator; Monte-Carlo credible intervals are calibrated for any sample size and naturally converge to Wald intervals as more data becomes available. We validate our approach empirically on benchmark datasets, demonstrating how our tests provide interpretable, statistically rigorous decisions under varying degrees of data availability and intersectionality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。