记忆部分数据可改善多分类公平性,关键在选对记忆样本
When Can Memorization Improve Fairness?
- 通过分析记忆数据的标签与群体分布,推导出偏差公式
- 确定完全消除三类公平性偏差所需记忆数据的最小比例
- 给出适用于模型公平性优化的研究者参考
我们研究在多分类问题中,通过记忆部分人口数据,能多大程度影响加性公平指标(统计均等、平等机会、等化机会)。给出了记忆导致偏差的显式表达式,该表达式基于记忆数据的标签与群体归属分布,以及分类器在未记忆数据上的偏差。同时,我们刻画了能消除三类指标偏差的记忆数据集特征,并提供了完全消除这些偏差所需的最小记忆概率质量上下界。
原文摘要 · Abstract (English)
We study to which extent additive fairness metrics (statistical parity, equal opportunity and equalized odds) can be influenced in a multi-class classification problem by memorizing a subset of the population. We give explicit expressions for the bias resulting from memorization in terms of the label and group membership distribution of the memorized dataset and the classifier bias on the unmemorized dataset. We also characterize the memorized datasets that eliminate the bias for all three metrics considered. Finally we provide upper and lower bounds on the total probability mass in the memorized dataset that is necessary for the complete elimination of these biases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。