arXiv:2505.23528cs.LG2025-05中稿 · IEEE Engineering i…被引 2

分析脑影像诊断阿尔茨海默病模型的公平性,找出年龄和种族偏见并提出缓解策略。

Comparative assessment of fairness definitions and bias mitigation strategies in machine learning-based diagnosis of Alzheimer's disease from MR images

  • 对比多种公平性定义与度量方法,筛选适合医疗场景的评估标准。
  • 针对种族和性别,拒绝选项分类使公平性提升46%~57%,年龄则用对抗去偏效果最好。
  • 提出新权衡指标,兼顾诊断准确率与公平性,适合临床应用。

本研究对基于磁共振成像(MRI)特征的轻度认知障碍(MCI)与阿尔茨海默病(AD)机器学习诊断模型进行了全面的公平性分析。在多队列数据集中,考察了年龄、种族和性别相关的偏差,以及编码这些敏感属性的代理特征的影响。评估了多种公平性定义与度量的可靠性,并基于最优度量,对比了常用的预处理、训练中和后处理去偏策略。此外,引入一种新的复合度量,综合考虑F1分数与等几率比(equalized odds ratio),以量化公平性与性能之间的权衡,适用于医疗诊断场景。结果表明,存在与年龄和种族相关的偏差,但性别偏差不显著。不同去偏策略在各类敏感属性上表现各异:对于种族和性别,拒绝选项分类分别使等几率比提升46%和57%,在MCI vs AD子问题中取得0.75和0.80的调和平均分;对于年龄,在同一子问题中,对抗去偏实现最高40%的等几率比改进,调和平均分为0.69。研究还揭示了与人口学特征相关的阿尔茨海默病神经病理变化及风险因素如何影响模型公平性。

原文摘要 · Abstract (English)

The present study performs a comprehensive fairness analysis of machine learning (ML) models for the diagnosis of Mild Cognitive Impairment (MCI) and Alzheimer's disease (AD) from MRI-derived neuroimaging features. Biases associated with age, race, and gender in a multi-cohort dataset, as well as the influence of proxy features encoding these sensitive attributes, are investigated. The reliability of various fairness definitions and metrics in the identification of such biases is also assessed. Based on the most appropriate fairness measures, a comparative analysis of widely used pre-processing, in-processing, and post-processing bias mitigation strategies is performed. Moreover, a novel composite measure is introduced to quantify the trade-off between fairness and performance by considering the F1-score and the equalized odds ratio, making it appropriate for medical diagnostic applications. The obtained results reveal the existence of biases related to age and race, while no significant gender bias is observed. The deployed mitigation strategies yield varying improvements in terms of fairness across the different sensitive attributes and studied subproblems. For race and gender, Reject Option Classification improves equalized odds by 46% and 57%, respectively, and achieves harmonic mean scores of 0.75 and 0.80 in the MCI versus AD subproblem, whereas for age, in the same subproblem, adversarial debiasing yields the highest equalized odds improvement of 40% with a harmonic mean score of 0.69. Insights are provided into how variations in AD neuropathology and risk factors, associated with demographic characteristics, influence model fairness.

公平性阿尔茨海默病MRI去偏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。