arXiv:2508.01460cs.CV2025-08被引 14

无需真实标签,用不确定性预测分割质量,提升医学影像分析可信度。

Uncertainty-Aware Segmentation Quality Prediction via Deep Learning Bayesian Modeling: Comprehensive Evaluation and Interpretation on Skin Cancer and Liver Segmentation

  • 结合不确定性和分割图,构建无真值的分割质量评估框架。
  • 在皮肤癌和肝脏分割数据集上,相关性达96.58,优于现有方法。
  • 融合多类不确定性指标,支持临床决策与模型可靠性解释。

图像分割是计算医学图像分析的关键步骤,通常通过Dice系数等指标在训练和验证阶段评估。然而,在缺乏人工标注的临床场景中,分割质量评估面临挑战,且缺乏可靠性指标的模型难以推广应用。为此,我们提出一种新框架,可在测试时无需真实标注即可预测分割质量。该框架包含两种互补方法:一种利用预测分割图与不确定性图,另一种整合原始输入图像、不确定性图与分割图。我们采用蒙特卡洛丢弃、集成学习和测试时增强对SwinUNet和基于ResNet50的特征金字塔网络进行贝叶斯改造,以量化不确定性。在2D皮肤病变和3D肝脏分割数据集上,评估了置信度图、熵、互信息及期望成对KL散度四种不确定性估计,分析其与分割质量指标的相关性。在HAM10000数据集上,框架达到R²=93.25,皮尔逊相关系数为96.58,显著优于先前方法;在3D肝脏分割中,测试时增强结合熵的方案获得R²=85.03,皮尔逊相关系数65.02,展现跨模态鲁棒性。此外,我们提出一种聚合策略,将多个不确定性估计融合为单张图像的综合评分,提供更稳健的质量评估。最后,通过Grad-CAM和UMAP嵌入分析,揭示了不确定性整合对模型行为与可靠性的关键影响。

原文摘要 · Abstract (English)

Image segmentation is a critical step in computational biomedical image analysis, typically evaluated using metrics like the Dice coefficient during training and validation. However, in clinical settings without manual annotations, assessing segmentation quality becomes challenging, and models lacking reliability indicators face adoption barriers. To address this gap, we propose a novel framework for predicting segmentation quality without requiring ground truth annotations during test time. Our approach introduces two complementary frameworks: one leveraging predicted segmentation and uncertainty maps, and another integrating the original input image, uncertainty maps, and predicted segmentation maps. We present Bayesian adaptations of two benchmark segmentation models-SwinUNet and Feature Pyramid Network with ResNet50-using Monte Carlo Dropout, Ensemble, and Test Time Augmentation to quantify uncertainty. We evaluate four uncertainty estimates: confidence map, entropy, mutual information, and expected pairwise Kullback-Leibler divergence on 2D skin lesion and 3D liver segmentation datasets, analyzing their correlation with segmentation quality metrics. Our framework achieves an R2 score of 93.25 and Pearson correlation of 96.58 on the HAM10000 dataset, outperforming previous segmentation quality assessment methods. For 3D liver segmentation, Test Time Augmentation with entropy achieves an R2 score of 85.03 and a Pearson correlation of 65.02, demonstrating cross-modality robustness. Additionally, we propose an aggregation strategy that combines multiple uncertainty estimates into a single score per image, offering a more robust and comprehensive assessment of segmentation quality. Finally, we use Grad-CAM and UMAP-based embedding analysis to interpret the model's behavior and reliability, highlighting the impact of uncertainty integration.

医学图像分割质量不确定性贝叶斯深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。