arXiv:2604.04858cs.LGq-bio.QM2026-04被引 2

工具包FairLogue可检测医疗模型中多重身份群体的公平性偏差。

FairLogue: A Toolkit for Intersectional Fairness Analysis in Clinical Machine Learning Models

  • 扩展公平性指标至交叉身份群体,支持观测与反事实分析
  • 发现交叉群体差异显著,如公平性差距达0.20,真阳性率差为0.33
  • 适合医疗AI开发者和伦理审查者用于识别深层不公平现象

算法公平性对医疗机器学习的公平可信至关重要。现有工具多聚焦单一维度比较,可能忽略交叉身份群体的复合不公。本文提出FairLogue,一个基于Python的工具包,包含三部分:1)观测框架,将人口均等、平等机会等指标扩展至交叉身份群体;2)基于治疗情境的反事实框架;3)针对交叉身份干预的广义反事实框架。在All of Us Controlled Tier V8数据集的青光眼手术预测任务中评估,使用逻辑回归,以种族和性别为受保护属性。观测分析显示模型性能中等(AUROC=0.709;准确率=0.651),但存在显著交叉公平性差距:人口均等差异达0.20,平等机会的真阳性率与假阳性率差距分别为0.33和0.15。反事实分析采用置换零分布,获得的不公平度量(u-value)接近零,表明在控制协变量后,观察到的不公可能源于随机性。FairLogue提供模块化方案,整合观测与反事实方法,实现临床机器学习中交叉偏差的量化评估。

原文摘要 · Abstract (English)

Objective: Algorithmic fairness is essential for equitable and trustworthy machine learning in healthcare. Most fairness tools emphasize single-axis demographic comparisons and may miss compounded disparities affecting intersectional populations. This study introduces Fairlogue, a toolkit designed to operationalize intersectional fairness assessment in observational and counterfactual contexts within clinical settings. Methods: Fairlogue is a Python-based toolkit composed of three components: 1) an observational framework extending demographic parity, equalized odds, and equal opportunity difference to intersectional populations; 2) a counterfactual framework evaluating fairness under treatment-based contexts; and 3) a generalized counterfactual framework assessing fairness under interventions on intersectional group membership. The toolkit was evaluated using electronic health record data from the All of Us Controlled Tier V8 dataset in a glaucoma surgery prediction task using logistic regression with race and gender as protected attributes. Results: Observational analysis identified substantial intersectional disparities despite moderate model performance (AUROC = 0.709; accuracy = 0.651). Intersectional evaluation revealed larger fairness gaps than single-axis analyses, including demographic parity differences of 0.20 and equalized odds true positive and false positive rate gaps of 0.33 and 0.15, respectively. Counterfactual analysis using permutation-based null distributions produced unfairness ("u-value") estimates near zero, suggesting observed disparities were consistent with chance after conditioning on covariates. Conclusion: Fairlogue provides a modular toolkit integrating observational and counterfactual methods for quantifying and evaluating intersectional bias in clinical machine learning workflows.

医疗AI公平性交叉公平反事实分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。