研究模型在不同群体间的公平与准确权衡,给出有限样本下最优逼近方法。
The Statistical Fairness-Accuracy Frontier
- 基于有限数据推导出最优估计器,适应不同群体的协变量分布
- 揭示样本量不足时各群体福利受不对称影响,提出最优采样分配策略
- 提供全前沿统一置信带,可量化不同权衡方案的可靠性
我们研究单一预测模型服务于多个群体时的公平性与准确性权衡问题。公平-准确(FA)帕累托前沿是理解这一权衡的有效工具,它刻画了无法在任一维度上改进而不损害另一维度的模型集合。尽管完全刻画该前沿需掌握数据分布的全部信息,本文聚焦有限样本情形,量化设计者在数据有限条件下对前沿任意点的逼近能力,并界定最坏情况下的差距。特别地,我们推导出依赖于设计者对协变量分布知识的最坏情况最优估计器。针对每个估计器,我们分析有限样本效应如何不对称地影响各群体福利,并识别最优样本分配策略。最后,我们提供了整个FA前沿的统一有限样本边界,得到可量化不同公平-准确权衡方案可靠性比较的置信带。
原文摘要 · Abstract (English)
We study fairness-accuracy tradeoffs when a single predictive model must serve multiple demographic groups. A useful tool for understanding this tradeoff is the fairness-accuracy (FA) Pareto frontier, which characterizes the set of models that cannot be improved in either fairness or accuracy without worsening the other. While characterizing the FA frontier requires full knowledge of the data distribution, we focus on the finite-sample regime, quantifying how well a designer can approximate any point on the frontier from limited data and bounding the worst-case gap. In particular, we derive worst-case-optimal estimators that depend on the designer's knowledge of the covariate distribution. For each estimator, we characterize how finite-sample effects asymmetrically impact each group's welfare and identify optimal sample allocation strategies. Finally, we provide uniform finite-sample bounds for the entire FA frontier, yielding confidence bands that quantify the reliability of welfare comparisons across alternative fairness-accuracy tradeoffs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。