让分类结果同时公平且最优,解决预测公平性难题
Fair Decisions from Calibrated Scores: Achieving Optimal Classification While Satisfying Sufficiency
- 基于校准得分设计随机化分类器,满足充分性公平约束
- 可实现最优分类性能,同时保持正向预测值与漏诊率的合理组合
- 适用于需兼顾公平与准确性的实际场景,如医疗、金融决策
基于预测概率(得分)的二分类是监督学习中的基础任务。虽然阈值化得分在无约束条件下是贝叶斯最优的,但使用单一阈值通常违反统计群体公平性要求。在独立性(统计均等)和分离性(等效机会)下,若得分已满足相应条件,阈值化即可;但该方法不适用于充分性:即使得分完全校准(包括真实类概率),阈值化后仍会破坏预测均等性。本文提出在有限组校准得分假设下,实现充分性约束下的最优二分类(随机化)的精确解。我们给出了可达的阳性预测值(PPV)与假阴性排除率(FOR)可行对的几何表征,并据此推导出仅依赖组校准得分与组成员身份的简单后处理算法,可实现最优分类器。由于充分性与分离性通常不可共存,我们进一步识别出在满足充分性前提下最接近分离性的分类器,并证明其亦可通过本算法获得,性能常接近最优。
原文摘要 · Abstract (English)
Binary classification based on predicted probabilities (scores) is a fundamental task in supervised machine learning. While thresholding scores is Bayes-optimal in the unconstrained setting, using a single threshold generally violates statistical group fairness constraints. Under independence (statistical parity) and separation (equalized odds), such thresholding suffices when the scores already satisfy the corresponding criterion. However, this does not extend to sufficiency: even perfectly group-calibrated scores -- including true class probabilities -- violate predictive parity after thresholding. In this work, we present an exact solution for optimal binary (randomized) classification under sufficiency, assuming finite sets of group-calibrated scores. We provide a geometric characterization of the feasible pairs of positive predictive value (PPV) and false omission rate (FOR) achievable by such classifiers, and use it to derive a simple post-processing algorithm that attains the optimal classifier using only group-calibrated scores and group membership. Finally, since sufficiency and separation are generally incompatible, we identify the classifier that minimizes deviation from separation subject to sufficiency, and show that it can also be obtained by our algorithm, often achieving performance comparable to the optimum.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。