ABROCA测偏差有陷阱,样本量和类别不均衡会夸大结果
ABROCA Distributions For Algorithmic Bias Assessment: Considerations Around Interpretation
- 分析ABROCA统计特性,发现其分布高度偏斜
- 在相似AUC下,偏斜导致误判为存在性能差异
- 类别不平衡时更需谨慎解读,适合做公平性评估的学者参考
算法偏差仍是学习分析中的核心问题。本文研究绝对介于ROC曲线下面积(ABROCA)度量的统计特性,该公平性指标通过ROC曲线的绝对差异量化群体间分类器表现差异。即使整体AUC相近,ABROCA仍能检测细微性能差异。我们对不同AUC差异和类别分布下的ABROCA进行了采样,发现其分布受样本量、AUC差异和类别不平衡影响,呈现显著偏斜。模拟结果显示,即便数据来自具有相同ROC曲线的人群,偏斜仍会导致ABROCA值被偶然放大。这表明,在评估分类器是否偏倚,尤其是类别不平衡场景下,必须谨慎解释ABROCA值。
原文摘要 · Abstract (English)
Algorithmic bias continues to be a key concern of learning analytics. We study the statistical properties of the Absolute Between-ROC Area (ABROCA) metric. This fairness measure quantifies group-level differences in classifier performance through the absolute difference in ROC curves. ABROCA is particularly useful for detecting nuanced performance differences even when overall Area Under the ROC Curve (AUC) values are similar. We sample ABROCA under various conditions, including varying AUC differences and class distributions. We find that ABROCA distributions exhibit high skewness dependent on sample sizes, AUC differences, and class imbalance. When assessing whether a classifier is biased, this skewness inflates ABROCA values by chance, even when data is drawn (by simulation) from populations with equivalent ROC curves. These findings suggest that ABROCA requires careful interpretation given its distributional properties, especially when used to assess the degree of bias and when classes are imbalanced.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。