arXiv:2505.08685cs.CV2025-05被引 10

多标注者评估提升医学影像分割模型可靠性

Calibration and Uncertainty for multiRater Volume Assessment in multiorgan Segmentation (CURVAS) challenge results

  • 用多标注者共识与分歧构建更全面的金标准
  • 最佳模型达高骰子系数且置信度估计准确
  • 适合关注临床可信度的医疗AI研究者

深度学习已成为医学图像分割的主流方法,但确保其可靠性和临床可用性仍面临标注变异、模型校准和不确定性估计等挑战。为此,我们发起多标注者体积评估的校准与不确定性(CURVAS)挑战,强调多标注者在建立更全面金标准中的关键作用,指出分割本质上具有主观性,需利用标注间差异来实现稳健模型评估。七支团队参与,提交多种深度学习模型,采用骰子相似系数(DSC)、期望校准误差(ECE)和连续概率评分(CRPS)等指标进行评估。通过引入共识与分歧金标准,分析模型对不确定性的处理能力及其置信度与真实性能的一致性。结果表明,良好校准的模型与高质量结果强相关;在多样化数据集上训练并融合预训练知识的模型表现出更强鲁棒性,尤其在偏离标准解剖结构时表现更优。最优模型同时获得高DSC值和良好校准的不确定性估计。本工作强调多标注者金标准、全面校准评估及不确定性感知评价对构建可信、可临床应用的深度学习分割模型的重要性。

原文摘要 · Abstract (English)

Deep learning (DL) has become the dominant approach for medical image segmentation, yet ensuring the reliability and clinical applicability of these models requires addressing key challenges such as annotation variability, calibration, and uncertainty estimation. This is why we created the Calibration and Uncertainty for multiRater Volume Assessment in multiorgan Segmentation (CURVAS), which highlights the critical role of multiple annotators in establishing a more comprehensive ground truth, emphasizing that segmentation is inherently subjective and that leveraging inter-annotator variability is essential for robust model evaluation. Seven teams participated in the challenge, submitting a variety of DL models evaluated using metrics such as Dice Similarity Coefficient (DSC), Expected Calibration Error (ECE), and Continuous Ranked Probability Score (CRPS). By incorporating consensus and dissensus ground truth, we assess how DL models handle uncertainty and whether their confidence estimates align with true segmentation performance. Our findings reinforce the importance of well-calibrated models, as better calibration is strongly correlated with the quality of the results. Furthermore, we demonstrate that segmentation models trained on diverse datasets and enriched with pre-trained knowledge exhibit greater robustness, particularly in cases deviating from standard anatomical structures. Notably, the best-performing models achieved high DSC and well-calibrated uncertainty estimates. This work underscores the need for multi-annotator ground truth, thorough calibration assessments, and uncertainty-aware evaluations to develop trustworthy and clinically reliable DL-based medical image segmentation models.

医学影像不确定性模型校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。