提出新后处理方法,提升多输出模型的公平性
Multi-Output Distributional Fairness via Post-Processing
- 用最优传输映射将不同群体输出移向经验沃尔什巴里中心
- 在多任务分类与表征学习中显著改善分布公平性
- 适用于多输出模型,可扩展至未见数据,计算高效
后处理方法因其直观性、低计算成本和良好可扩展性,正成为提升机器学习模型公平性的关键手段。然而,现有方法大多针对特定任务的公平性度量,且仅限于单输出模型。本文提出一种适用于多输出模型(如多任务/多分类分类和表征学习)的后处理方法,以增强模型的分布平等性——一种任务无关的公平性度量。以往实现分布平等的方法依赖于模型输出的(逆)累积分布函数,限制了其在多输出场景的应用。本文通过引入最优传输映射,将不同群体的模型输出向其经验沃尔什巴里中心移动。为降低精确巴里中心计算复杂度,提出近似技术;并设计核回归方法,将该过程推广至未见数据。实验在多任务/多分类分类与表征学习任务上对比多种基线,验证了所提方法的有效性。
原文摘要 · Abstract (English)
The post-processing approaches are becoming prominent techniques to enhance machine learning models' fairness because of their intuitiveness, low computational cost, and excellent scalability. However, most existing post-processing methods are designed for task-specific fairness measures and are limited to single-output models. In this paper, we introduce a post-processing method for multi-output models, such as the ones used for multi-task/multi-class classification and representation learning, to enhance a model's distributional parity, a task-agnostic fairness measure. Existing methods for achieving distributional parity rely on the (inverse) cumulative density function of a model's output, restricting their applicability to single-output models. Extending previous works, we propose to employ optimal transport mappings to move a model's outputs across different groups towards their empirical Wasserstein barycenter. An approximation technique is applied to reduce the complexity of computing the exact barycenter and a kernel regression method is proposed to extend this process to out-of-sample data. Our empirical studies evaluate the proposed approach against various baselines on multi-task/multi-class classification and representation learning tasks, demonstrating the effectiveness of the proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。