arXiv:2606.00656cs.LGcs.AI2026-06中稿 · ICML

提出多分类公平分类最优解法,兼顾准确率与公平性。

Demystifying the Optimal Fair Classifier in Multi-Class Classification

  • 构建可解析的公平约束下最优分类器模型
  • 在多个数据集上实现准确率与公平性的更好平衡
  • 适合关注多分类公平性的研究者与从业者

在多分类任务中确保不同群体的公平对待面临巨大挑战,因机器学习模型固有的偏差难以消除。现有多数偏见缓解方法仅适用于二分类场景,而多维输出与复杂公平机制使得其向多分类推广既不直接也无效。本文聚焦两个未解核心问题:(i) 揭示多分类下准确率-公平性前沿的理论特征;(ii) 设计可在不同训练阶段达到该前沿的实用算法。我们首先给出一个可解析的概率化最优分类器形式。基于此,提出两种无属性依赖的算法:一种训练中干预的内处理方法(通过约简法),一种输出概率后处理微调的插件估计法。理论分析表明,两者均收敛至最优准确率-公平性帕累托前沿。多数据集实验验证了所提方法在平衡准确率与公平性方面的优越性能。

原文摘要 · Abstract (English)

Ensuring fair and equitable treatment across diverse groups, particularly in multi-class classification tasks, poses a significant challenge due to the persistent biases inherent in machine learning models. Most existing bias mitigation techniques are tailored to binary settings, and the presence of multi-dimensional outputs and complex fairness mechanisms makes their extension to multi-class scenarios neither straightforward nor effective. In this paper, we investigate two fundamental, unresolved challenges in fair classification: (i) characterizing the optimal accuracy-fairness frontier in multi-class settings, and (ii) designing practical algorithms that attain this optimum in different training phases. To tackle these challenges, we first specify an analytically tractable probabilistic formulation of the optimal classifier under fairness constraints. Building upon this, we propose two attribute-blind algorithms to enforce fairness requirements in practice: an in-processing approach for fairness intervention during training via the reduction approach, and a post-processing approach for fine-tuning output probabilities with plug-in estimation. Theoretical analysis reveals that both methods converge to the optimal accuracy-fairness Pareto frontier. Experiments conducted on multiple datasets demonstrate the superior performance of our methods in balancing accuracy and fairness.

公平分类多分类优化算法公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。