通过分解特征空间,让模型在公平与准确间更好平衡。
Understanding Fairness and Prediction Error through Subspace Decomposition and Influence Analysis
- 将特征空间拆分为目标相关、敏感属性和共享三部分
- 移除敏感信息后,公平性提升且预测误差变化可控
- 适合关注模型公平性但不想牺牲性能的研究者
机器学习模型虽广泛应用,却常继承并放大历史偏见,导致不公平结果。传统公平方法多在预测层面施加约束,未触及数据表示中的深层偏见。本文提出一种理论框架,通过充分维数缩减将特征空间分解为目标相关、敏感属性与共享分量,通过选择性移除敏感信息来调节公平性与预测效用的权衡。我们分析了随着共享子空间增加,预测误差与公平差距的演化规律,并利用影响函数量化其对参数估计渐近行为的影响。在合成与真实数据集上的实验验证了理论发现,表明该方法能有效提升公平性同时保持预测性能。
原文摘要 · Abstract (English)
Machine learning models have achieved widespread success but often inherit and amplify historical biases, resulting in unfair outcomes. Traditional fairness methods typically impose constraints at the prediction level, without addressing underlying biases in data representations. In this work, we propose a principled framework that adjusts data representations to balance predictive utility and fairness. Using sufficient dimension reduction, we decompose the feature space into target-relevant, sensitive, and shared components, and control the fairness-utility trade-off by selectively removing sensitive information. We provide a theoretical analysis of how prediction error and fairness gaps evolve as shared subspaces are added, and employ influence functions to quantify their effects on the asymptotic behavior of parameter estimates. Experiments on both synthetic and real-world datasets validate our theoretical insights and show that the proposed method effectively improves fairness while preserving predictive performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。