纠正数据分布偏差后,才能真实评估公平性代价。
Fix Representation (Optimally) Before Fairness: Finite-Sample Shrinkage Population Correction and the True Price of Fairness Under Subpopulation Shift
- 提出有限样本下最优的收缩重加权方法修正子群比例偏差
- 实验证明,公平性提升看似提高准确率,实为基线未校准所致
- 建议先优化表示再评估公平性,避免虚假权衡
机器学习实践中常出现预测准确率与群体公平性之间的矛盾,但有时公平性干预反而提升准确率。我们发现这两种现象都可能是训练数据未能反映真实子群比例所致。在子群分布稳定但群体比例偏移(subpopulation shift)条件下,我们证明:(i) 全量重要性权重修正虽渐近无偏,但在有限样本下非最优;(ii) 最优有限样本修正为一种介于目标与训练混合分布之间的收缩重加权;(iii) 表面的“公平性提升准确率”可能源于与未正确加权的基线对比。我们提出可操作的评估协议:先最优地纠正表示,再比较公平性干预,以分离出公平性的真正不可减少代价。在合成数据及真实世界基准(Adult、COMPAS)上的实验验证了理论预测,表明该协议能消除虚假权衡,揭示真实的公平性-效用边界。
原文摘要 · Abstract (English)
Machine learning practitioners frequently observe tension between predictive accuracy and group fairness constraints -- yet sometimes fairness interventions appear to improve accuracy. We show that both phenomena can be artifacts of training data that misrepresents subgroup proportions. Under subpopulation shift (stable within-group distributions, shifted group proportions), we establish: (i) full importance-weighted correction is asymptotically unbiased but finite-sample suboptimal; (ii) the optimal finite-sample correction is a shrinkage reweighting that interpolates between target and training mixtures; (iii) apparent "fairness helps accuracy" can arise from comparing fairness methods to an improperly-weighted baseline. We provide an actionable evaluation protocol: fix representation (optimally) before fairness -- compare fairness interventions against a shrinkage-corrected baseline to isolate the true, irreducible price of fairness. Experiments on synthetic and real-world benchmarks (Adult, COMPAS) validate our theoretical predictions and demonstrate that this protocol eliminates spurious tradeoffs, revealing the genuine fairness-utility frontier.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。