通过肺部裁剪减少胸部X光中的种族偏见,提升模型公平性。
The Impact of Preprocessing Methods on Racial Encoding and Model Robustness in CXR Diagnosis
- 用肺部边界框裁剪图像,抑制种族相关干扰信号。
- 裁剪后模型识别种族的准确率下降60%以上,诊断性能基本不变。
- 为临床AI提供无需牺牲精度的公平性优化方案,适合医疗影像开发者。
深度学习模型能以高准确率从胸部X光(CXR)中识别种族身份,引发对种族捷径学习的广泛担忧——模型可能无意间根据种族身份系统性地偏倚诊断结果。这类偏见威胁医疗公平与模型可靠性,可能导致特定群体被持续误诊。由于种族捷径信号具有非局部、分布式的特性,图像预处理方法可能影响其学习过程,但该潜力尚未充分探索。本文研究了肺部掩码、肺部裁剪和对比度受限自适应直方图均衡化(CLAHE)等预处理方法的效果。这些方法旨在抑制编码种族信息的虚假线索,同时保留诊断准确性。实验表明,简单的基于边界框的肺部裁剪可有效降低种族捷径学习,且在多项标准评估中保持诊断模型性能,规避了常被提及的公平性-准确性权衡问题。
原文摘要 · Abstract (English)
Deep learning models can identify racial identity with high accuracy from chest X-ray (CXR) recordings. Thus, there is widespread concern about the potential for racial shortcut learning, where a model inadvertently learns to systematically bias its diagnostic predictions as a function of racial identity. Such racial biases threaten healthcare equity and model reliability, as models may systematically misdiagnose certain demographic groups. Since racial shortcuts are diffuse - non-localized and distributed throughout the whole CXR recording - image preprocessing methods may influence racial shortcut learning, yet the potential of such methods for reducing biases remains underexplored. Here, we investigate the effects of image preprocessing methods including lung masking, lung cropping, and Contrast Limited Adaptive Histogram Equalization (CLAHE). These approaches aim to suppress spurious cues encoding racial information while preserving diagnostic accuracy. Our experiments reveal that simple bounding box-based lung cropping can be an effective strategy for reducing racial shortcut learning while maintaining diagnostic model performance, bypassing frequently postulated fairness-accuracy trade-offs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。