arXiv:2606.19300cs.CVcs.LG2026-06中稿 · MIUA2016

MC Dropout虽能识别错误,但无法发现关键区域的置信度失准。

Confidence is Not Reliability: Rethinking MC Dropout in Brain Tumour Segmentation

论文配图:Confidence is Not Reliability: Rethinking MC Dropout in Brain Tumour Segmentation
图 1 · 摘自论文原文
  • 用蒙特卡洛丢弃法评估像素级不确定性,排序错误区域效果好
  • 高不确定性患者组肿瘤分割精度显著下降(中位Dice 0.835 vs 0.925)
  • 需额外检查关键子区域校准性,仅看整体指标有风险

多参数MRI中的胶质瘤分割对治疗规划至关重要。分割模型在治疗关键区域无声失效构成患者安全风险,而基于重叠的指标如Dice分数难以暴露此问题。本文通过126例BraTS21数据集上的双模型实证研究,评估了预训练SegResNet与局部训练的UNet-Res在蒙特卡洛丢弃(MC Dropout)下的表现。结果显示,尽管保持分割准确率(|ΔDice| < 0.01),熵值的不确定性-误差对齐度极高(AUROC ≈ 0.97),能有效区分错误与正确像素;基于熵的患者分层发现高不确定性组整体肿瘤Dice中位数为0.835,显著低于低不确定性组的0.925,表明不确定性可作临床分诊信号。然而全局对齐掩盖了区域差异:尽管两模型整体AUROC相近,但UNet-Res在增强肿瘤区熵值极低(0.054),预期校准误差(ECE)高达0.915,且该区域的Dice仅为0.714,显示其在最临床关键区域严重校准失败,这种缺陷仅靠标准的Dice和AUROC无法察觉。结论:强不确定性-误差对齐是必要条件,但不足以保障临床安全,必须结合子区域校准评估来选择部署模型。

原文摘要 · Abstract (English)

Glioma segmentation in multiparametric MRI is a critical component of treatment planning. A segmentation model that fails silently on treatment-critical sub-regions represents a patient safety risk that overlap-based metrics such as Dice scores cannot expose. We ask whether voxel-level uncertainty estimation via Monte Carlo (MC) Dropout can reliably identify segmentation errors in clinically critical sub-regions, and whether calibration failure modes are detectable from standard reporting metrics alone. In an empirical two-model case study on 126 BraTS21 patients, we evaluate a high-performance pretrained SegResNet and a locally trained UNet with residual units (UNet-Res). MC dropout preserved segmentation accuracy ($|Δ\text{Dice}|$ $<0.01$) while achieving strong uncertainty-error alignment (AUROC for entropy (H) $\approx$0.97), indicating uncertainty correctly ranks erroneous voxels above correct ones. Entropy-based patient stratification identified a high-uncertainty subgroup with substantially lower segmentation performance (median whole-tumour Dice $0.835$ vs. $0.925$), supporting uncertainty as a practical triage signal. However, global alignment can mask important region-specific differences. Despite similar AUROC, UNet-Res exhibited near-zero enhancing tumour entropy ($0.054$) and Expected Calibration Error (ECE) of $0.915$, with a Dice of only $0.714$, indicating severely miscalibrated confidence on the most clinically critical sub-region, a failure mode invisible to standard Dice and AUROC reporting. These findings demonstrate that strong uncertainty-error alignment is necessary but insufficient for clinical safety: sub-region-specific calibration assessment must accompany AUROC evaluation when selecting models for clinical deployment.

医学图像不确定性模型校准分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。