用置信度方法提升脑胶质瘤分割可靠性,让模型自知不确定
CONSeg: Voxelwise Glioma Conformal Segmentation
- 基于置信区间理论,为分割结果生成不确定性评分
- 不确定性越高的区域,分割准确率显著下降(p<0.001)
- 适合临床医生需要判断模型可信度的场景
脑胶质瘤分割对临床决策至关重要。本文将置信性预测(Conformal Prediction, CP)应用于脑胶质瘤分割,以量化模型不确定性。使用UCSF和UPenn数据集,分别划分训练(70%)、验证(10%)、校准(10%)、测试(10%)及外部测试(70%)集。采用UNet模型,通过预测归一化确定最优阈值0.5。基于内部/外部校准集的非符合度得分选择置信阈值,对内部与外部测试集应用CP,并报告覆盖率。定义不确定性比率(UR),分析其与Dice相似系数(DSC)的相关性。根据UR将病例分为确定与不确定组,比较两组DSC差异。同时评估了BraTS融合模型(BFMS)中UR与DSC的相关性。结果显示,基础模型在内部与外部测试集上的DSC分别为0.8628和0.8257,内部与外部测试集的CP覆盖率为0.9982和0.9977。统计分析表明,所有测试集上UR与DSC呈显著负相关(p<0.001),且在BFMS中,高不确定性也对应更低的DSC(p<0.001)。确定组的DSC显著高于不确定组(p<0.001)。结论:置信性预测能有效量化脑胶质瘤分割中的不确定性,提升模型可靠性并增强人机协作。
原文摘要 · Abstract (English)
Background and Purpose: Glioma segmentation is crucial for clinical decisions and treatment planning. Uncertainty quantification methods, including conformal prediction (CP), can enhance segmentation models reliability. This study aims to use CP in glioma segmentation. Methods: We used the UCSF and UPenn glioma datasets, with the UCSF dataset split into training (70%), validation (10%), calibration (10%), and test (10%) sets, and the UPenn dataset divided into external calibration (30%) and external test (70%) sets. A UNet model was trained, and its optimal threshold was set to 0.5 using prediction normalization. To apply CP, the conformal threshold was selected based on the internal/external calibration nonconformity score, and CP was subsequently applied to the internal/external test sets, with coverage reported for all. We defined the uncertainty ratio (UR) and assessed its correlation with the Dice score coefficient (DSC). Additionally, we categorized cases into certain and uncertain groups based on UR and compared their DSC. We also evaluate the correlation between UR and DSC of the BraTS fusion model segmentation (BFMS), and compare DSC in the certain and uncertain subgroups. Results: The base model achieved a DSC of 0.8628 and 0.8257 on the internal and external test sets, respectively. The CP coverage was 0.9982 for the internal test set and 0.9977 for the external test set. Statistical analysis showed a significant negative correlation between UR and DSC for test sets (p<0.001). UR was also linked to significantly lower DSCs in the BFMS (p<0.001). Additionally, certain cases had significantly higher DSCs than uncertain cases in test sets and the BFMS (p<0.001). Conclusion: CP effectively quantifies uncertainty in glioma segmentation. Using CONSeg improves the reliability of segmentation models and enhances human-computer interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。