arXiv:2504.16238cs.LG2025-04

提出可通用的公平性后处理框架,不改模型也能调公平性。

General Post-Processing Framework for Fairness Adjustment of Machine Learning Models

  • 把训练时的公平性方法挪到事后调整,不改训练流程。
  • 在真实数据集上达成与对抗去偏相当的公平性-精度平衡。
  • 适合想快速提升模型公平性的开发者和合规团队。

随着机器学习在信贷评估、公共政策和人才招聘等关键领域广泛应用,确保公平性不仅是法律要求,也是道德责任。本文提出一种新型公平性调整框架,适用于回归与分类等多种任务,并支持多种公平性度量。不同于传统的预处理、训练中或事后处理方法,本方法将训练阶段的公平性技术改造为事后处理步骤。通过将公平性调整与模型训练过程解耦,该框架在保持模型平均性能的同时,提升了开发灵活性。其核心优势包括:无需定制损失函数、可用不同数据集进行公平性调节、支持作为黑盒系统的专有模型,以及提供可解释的公平性调整分析。我们在真实数据集上对比了对抗去偏方法,验证了该框架在公平性与准确性权衡上的有效性。

原文摘要 · Abstract (English)

As machine learning increasingly influences critical domains such as credit underwriting, public policy, and talent acquisition, ensuring compliance with fairness constraints is both a legal and ethical imperative. This paper introduces a novel framework for fairness adjustments that applies to diverse machine learning tasks, including regression and classification, and accommodates a wide range of fairness metrics. Unlike traditional approaches categorized as pre-processing, in-processing, or post-processing, our method adapts in-processing techniques for use as a post-processing step. By decoupling fairness adjustments from the model training process, our framework preserves model performance on average while enabling greater flexibility in model development. Key advantages include eliminating the need for custom loss functions, enabling fairness tuning using different datasets, accommodating proprietary models as black-box systems, and providing interpretable insights into the fairness adjustments. We demonstrate the effectiveness of this approach by comparing it to Adversarial Debiasing, showing that our framework achieves a comparable fairness/accuracy tradeoff on real-world datasets.

公平性后处理可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。