用多维度交叉分析揭示临床模型中的隐性偏见,发现多数差异实为随机所致。
Evaluating Intersectional Fairness across Clinical Machine Learning Use Cases using Fairlogue and the All of Us Research Program
- 采用交叉公平性审计工具,同时评估种族与性别等多重身份组合
- 发现交叉分析揭示的不公平程度远超单一属性分析
- 通过反事实检验表明多数差异可由随机分组解释,非系统性偏见
医疗数据中的交叉性偏见可能导致临床机器学习模型产生复合性不公,但当前大多数公平性评估仍独立处理各人口属性。本文使用 FairLogue 工具包,在 All of Us 研究项目数据上对多个临床预测任务进行交叉公平性审计。选取两个已发表模型:(A) 选择性5-羟色胺再摄取抑制剂相关出血事件预测;(B) 心房颤动患者两年内中风风险预测。在种族、性别及交叉子群体上计算观察性公平性指标,并通过反事实分析判断不公是否源于群体归属。结果显示,交叉分析揭示的差距显著大于单轴分析;但反事实诊断表明,多数观测到的差距与随机分组下的预期水平相当。研究强调交叉公平性审计的重要性,并展示 FairLogue 如何深化对临床机器学习系统偏见的理解。
原文摘要 · Abstract (English)
Intersectional biases in healthcare data can produce compound disparities in clinical machine learning models, yet most fairness evaluations assess demographic attributes independently. FairLogue, a toolkit for intersectional fairness auditing, was applied across multiple clinical prediction tasks to evaluate disparities across combined demographic groups. Using the All of Us dataset, two published models were selected for replication and evaluation: (A) prediction of selective serotonin reuptake inhibitor associated bleeding events and (B) two-year stroke risk in patients with atrial fibrillation. Observational fairness metrics were computed across race, gender, and intersectional subgroups, followed by counterfactual analysis to evaluate whether disparities were attributable to group membership. Intersectional evaluation revealed larger disparities than single-axis analyses; however, counterfactual diagnostics indicated that most observed disparities were comparable to those expected under randomized group membership. These results highlight the importance of intersectional fairness auditing and demonstrate how FairLogue provides deeper insight into bias in clinical machine learning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。