Dice系数天然偏向男性,小器官分割误差会因性别不同被不公平低估
Sex-based Bias Inherent in the Dice Similarity Coefficient: A Model Independent Analysis for Multiple Anatomical Structures
- 用统一合成误差测试50人MRI标注,剥离模型影响
- 小器官平均DSC差异达0.03,中等器官0.01,大器官几乎无差异
- 提醒研究者:评估公平性时需警惕指标本身带来的性别偏差
基于重叠的评价指标(如Dice相似系数,DSC)对较小结构的分割误差惩罚更重。由于器官大小存在性别差异,相同程度的分割误差在女性中可能导致更低的DSC值。尽管已有研究关注模型或数据集中的性别差异,但尚无工作分析DSC本身可能引入的偏倚。本研究在理想化设置下,独立于具体模型,量化了DSC与归一化DSC的性别差异。我们对50名受试者的手动MRI标注施加等量合成误差以确保性别可比性。即使微小误差(如1 mm边界偏移)也导致系统性性别差异。小结构平均DSC差异约0.03,中等结构约0.01;仅大结构(肺、肝)基本不受影响,性别差异接近零。结果表明,使用DSC作为评估指标时,不应期望男女得分相同,因为该指标本身即具偏倚。一个在误差幅度上表现相当的模型,其观察到的DSC值仍可能显示性别差异。本研究揭示了一个此前未被充分关注的来源——评价指标本身造成的性别差异,而非模型行为所致。认识到这一点对于医疗图像分析中的准确、公平评估至关重要。
原文摘要 · Abstract (English)
Overlap-based metrics such as the Dice Similarity Coefficient (DSC) penalize segmentation errors more heavily in smaller structures. As organ size differs by sex, this implies that a segmentation error of equal magnitude may result in lower DSCs in women due to their smaller average organ volumes compared to men. While previous work has examined sex-based differences in models or datasets, no study has yet investigated the potential bias introduced by the DSC itself. This study quantifies sex-based differences of the DSC and the normalized DSC in an idealized setting independent of specific models. We applied equally-sized synthetic errors to manual MRI annotations from 50 participants to ensure sex-based comparability. Even minimal errors (e.g., a 1 mm boundary shift) produced systematic DSC differences between sexes. For small structures, average DSC differences were around 0.03; for medium-sized structures around 0.01. Only large structures (i.e., lungs and liver) were mostly unaffected, with sex-based DSC differences close to zero. These findings underline that fairness studies using the DSC as an evaluation metric should not expect identical scores between men and women, as the metric itself introduces bias. A segmentation model may perform equally well across sexes in terms of error magnitude, even if observed DSC values suggest otherwise. Importantly, our work raises awareness of a previously underexplored source of sex-based differences in segmentation performance. One that arises not from model behavior, but from the metric itself. Recognizing this factor is essential for more accurate and fair evaluations in medical image analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。