对比深度与传统方法在脑部MRI分割中的种族性别偏见
Investigating Demographic Bias in Brain MRI Segmentation: A Comparative Study of Deep-Learning and Non-Deep-Learning Methods
- 用四种人群数据测试四种分割模型,评估种族性别影响
- 训练数据与测试对象种族匹配时,部分模型精度显著提升
- 仅一个模型保留种族对核团体积的影响,其余均消失
深度学习分割算法虽推动了医学影像分析发展,但数据内在偏见引发公平性担忧。本文评估了UNesT、nnU-Net、CoTr三种深度学习模型及传统基于图谱的ANTs方法,在黑人女性、黑人男性、白人女性、白人男性四类人群中分割左右伏隔核(NAc)的表现。使用人工标注金标准数据进行训练与测试。研究分两部分:第一部分评估模型分割性能,第二部分分析生成体积受种族、性别及其交互作用的影响。采用量化公平性指标衡量分割公平性,并通过线性混合模型分析人口学变量对准确率和体积的影响。结果显示,针对同种族测试样本训练时,ANTs和UNesT的分割精度显著提高,而nnU-Net表现稳定不受匹配影响。手动标注显示性别差异存在于体积中,该效应在多数模型中仍可复现;但种族差异仅在一个模型中存在,其余均消失。
原文摘要 · Abstract (English)
Deep-learning-based segmentation algorithms have substantially advanced the field of medical image analysis, particularly in structural delineations in MRIs. However, an important consideration is the intrinsic bias in the data. Concerns about unfairness, such as performance disparities based on sensitive attributes like race and sex, are increasingly urgent. In this work, we evaluate the results of three different segmentation models (UNesT, nnU-Net, and CoTr) and a traditional atlas-based method (ANTs), applied to segment the left and right nucleus accumbens (NAc) in MRI images. We utilize a dataset including four demographic subgroups: black female, black male, white female, and white male. We employ manually labeled gold-standard segmentations to train and test segmentation models. This study consists of two parts: the first assesses the segmentation performance of models, while the second measures the volumes they produce to evaluate the effects of race, sex, and their interaction. Fairness is quantitatively measured using a metric designed to quantify fairness in segmentation performance. Additionally, linear mixed models analyze the impact of demographic variables on segmentation accuracy and derived volumes. Training on the same race as the test subjects leads to significantly better segmentation accuracy for some models. ANTs and UNesT show notable improvements in segmentation accuracy when trained and tested on race-matched data, unlike nnU-Net, which demonstrates robust performance independent of demographic matching. Finally, we examine sex and race effects on the volume of the NAc using segmentations from the manual rater and from our biased models. Results reveal that the sex effects observed with manual segmentation can also be observed with biased models, whereas the race effects disappear in all but one model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。