arXiv:2511.00477eess.IVcs.AI2025-11被引 5

发现乳腺癌分割模型对年轻患者偏差大,根源是图像本身难学而非标注差。

Investigating Label Bias and Representational Sources of Age-Related Disparities in Medical Segmentation

  • 通过控制实验排除标注质量影响,发现年轻患者图像本就更难分割。
  • 验证了'有偏尺子效应':错误标注会扭曲模型真实偏差表现。
  • 强调公平需关注图像特征差异,而非简单平衡数据量。

医学影像中的算法偏见可能加剧健康不平等,但其在分割任务中的成因仍不清晰。尽管分类任务中公平性研究较多,分割因临床重要性而亟待关注。在乳腺癌分割中,模型对年轻患者的性能显著下降,传统认为源于乳腺密度的生理差异。我们审计了MAMA-MIA数据集,量化了年龄相关的标签偏差,并揭示了关键的'有偏尺子效应':验证集中系统性错误的标注会误导模型实际偏差的评估。然而,该偏差是否源于标注质量或图像本身的挑战性尚不明确。通过控制实验,我们系统排除了标注质量敏感性和病例难度失衡的假设。按难度平衡训练数据无法缓解差距,表明年轻患者病例内在更难学习。我们提供了直接证据:在有偏的机器生成标注上训练会学习并放大系统性偏差,这对自动化标注流程具有重要意义。本研究提出诊断医疗分割算法偏见的系统框架,证明实现公平需解决质性分布差异,而非仅平衡样本数量。

原文摘要 · Abstract (English)

Algorithmic bias in medical imaging can perpetuate health disparities, yet its causes remain poorly understood in segmentation tasks. While fairness has been extensively studied in classification, segmentation remains underexplored despite its clinical importance. In breast cancer segmentation, models exhibit significant performance disparities against younger patients, commonly attributed to physiological differences in breast density. We audit the MAMA-MIA dataset, establishing a quantitative baseline of age-related bias in its automated labels, and reveal a critical Biased Ruler effect where systematically flawed labels for validation misrepresent a model's actual bias. However, whether this bias originates from lower-quality annotations (label bias) or from fundamentally more challenging image characteristics remains unclear. Through controlled experiments, we systematically refute hypotheses that the bias stems from label quality sensitivity or quantitative case difficulty imbalance. Balancing training data by difficulty fails to mitigate the disparity, revealing that younger patient cases are intrinsically harder to learn. We provide direct evidence that systemic bias is learned and amplified when training on biased, machine-generated labels, a critical finding for automated annotation pipelines. This work introduces a systematic framework for diagnosing algorithmic bias in medical segmentation and demonstrates that achieving fairness requires addressing qualitative distributional differences rather than merely balancing case counts.

医疗影像算法公平性分割偏差标注质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。