arXiv:2607.08953cs.LG2026-07

系统评估多阶段公平性方法组合,发现联合策略更有效但效果依赖场景。

FairSelect: A Systematic Evaluation of Multi-Level and Intersectional Algorithmic Fairness

  • 构建多阶段公平性评估框架,支持预处理、内处理、后处理方法组合测试。
  • 合成数据中联合策略平均提升公平性,真实临床数据中部分组合同时提升性能与公平性。
  • 适用于医疗机器学习中的公平性策略选择,尤其关注交叉群体差异的场景。

算法公平性方法被广泛用于识别和缓解机器学习模型中的偏见,但现有评估大多孤立进行,且仅沿单一人口属性轴展开。这限制了在实际中选择公平性策略的指导性,因为偏见可能出现在交叉子群体及建模生命周期多个阶段。本文提出 FairSelect,一个系统化评估多阶段(预处理、内处理、后处理)公平性缓解策略的方法框架。该工具支持多种模型架构、交叉子群体评估,并可比较基线、单方法与多层级配置下的公平性-效用权衡。通过设计具有特定偏见机制的合成临床数据集以及对心房颤动患者两年卒中风险预测的真实世界复现进行验证。合成实验表明,针对性公平方法通常能降低目标子群体偏差,而组合策略带来更大的平均公平性改进,仅伴随适度效用损失。在临床预测任务中,干预效果高度可变:某些组合同时提升公平性与预测性能,而其他组合则无效甚至产生反效果。结果表明,公平性干预存在非加性和情境依赖性交互作用。FairSelect 提供了一套实用框架,可在保持模型性能的前提下系统识别改善子群体公平性的策略。

原文摘要 · Abstract (English)

Algorithmic fairness methods are increasingly used to identify and mitigate bias in machine learning models, yet most approaches are evaluated in isolation and along single demographic axes. This limits practical guidance for selecting fairness strategies, where disparities may arise across intersectional subgroups and across multiple stages of the modeling lifecycle. This work presents FairSelect, a toolkit for systematically evaluating fairness mitigation strategies applied individually and in combination across preprocessing, inprocessing, and postprocessing stages. FairSelect supports multiple model architectures, intersectional subgroup evaluation, and comparison of fairness utility tradeoffs across baseline, single method, and multi level configurations. The framework was validated using synthetic clinical datasets designed to represent specific bias mechanisms and a real-world replication of two-year stroke risk prediction among patients with atrial fibrillation. Synthetic experiments showed that targeted fairness methods generally reduced intended subgroup disparities, while combined strategies produced larger average fairness improvements with modest utility tradeoffs. In the clinical prediction task, mitigation effects were highly variable, with some combinations improving both fairness and predictive performance while others were ineffective or counterproductive. These findings demonstrate that fairness interventions interact in nonadditive and context dependent ways. FairSelect provides a practical framework for systematically identifying fairness strategies that improve subgroup equity while preserving model performance in clinical machine learning.

算法公平医疗AI交叉公平多阶段评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。