让分类结果更公平:在多类场景下约束集合预测的偏差与大小。
Set to Be Fair: Demographic Parity Constraints for Set-Valued Classification
- 提出两种公平约束下的集合分类方法,兼顾准确性与公平性。
- 理论证明方法在无分布假设下能有效控制偏差和大小误差。
- 适合关注算法公平性的机器学习研究者与应用开发者。
集合值分类用于多类场景,当类别间易混淆时可避免误导性预测。但其应用可能放大歧视性偏见,因此需在公平性约束下设计方法。本文研究在人口统计均等性(demographic parity)和期望大小约束下的集合值分类问题。提出两种互补策略:基于理想情况的最小化风险方法,以及计算高效的代理方法,后者优先保障约束满足。针对两者,推导出最优公平集合分类器的闭式表达,并据此构建可直接使用的数据驱动预测程序。建立了两种方法在无分布假设下对大小与公平性约束违反的收敛速率;在弱假设下,还给出了理想方法的额外风险上界。实验表明两种策略均有效,且代理方法效率更高。
原文摘要 · Abstract (English)
Set-valued classification is used in multiclass settings where confusion between classes can occur and lead to misleading predictions. However, its application may amplify discriminatory bias motivating the development of set-valued approaches under fairness constraints. In this paper, we address the problem of set-valued classification under demographic parity and expected size constraints. We propose two complementary strategies: an oracle-based method that minimizes classification risk while satisfying both constraints, and a computationally efficient proxy that prioritizes constraint satisfaction. For both strategies, we derive closed-form expressions for the (optimal) fair set-valued classifiers and use these to build plug-in, data-driven procedures for empirical predictions. We establish distribution-free convergence rates for violations of the size and fairness constraints for both methods, and under mild assumptions we also provide excess-risk bounds for the oracle-based approach. Empirical results demonstrate the effectiveness of both strategies and highlight the efficiency of our proxy method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。