对比深度集成与单模型在脑肿瘤分割中的可靠性,发现集成更能识别出错情况。
Reliability analysis for BraTS-GoAT segmentation: a controlled robustness study of deep-ensemble uncertainty
- 用3种子模型深度集成,通过成员间分歧检测错误
- 在合成数据扰动下,集成分歧上升25%-33%,单模型置信度却不变
- 适合关注模型可靠性、临床部署风险的开发者和医生
深度网络在分布内脑肿瘤分割表现准确,但在输入偏离训练数据时可能无声失效。这正是BraTS-GoAT泛化性任务的核心挑战。我们不仅评估分割性能,更关注模型不确定性是否能识别自身错误。在BraTS-GoAT(任务3)上,训练了5折交叉验证的nnU-Net基线(每例一个预测)和3种子深度集成。两者均在区域相关掩码上评估校准性和错误检测,按病例聚合。分布内,3种子集成在相同保留划分上小幅优于强基线模型,尤其在校准性方面提升明显。分布外差异显现:使用分级合成伪影作为扫描差异的代理,在控制鲁棒性研究中,单模型置信度保持平稳,而其准确率和校准性下降;成员间分歧则急剧上升,比干净条件下高出约1/4至1/3,是单模型响应的数倍。官方验证榜单上,5折集成获得全肿瘤Dice为0.87。泛化差距集中在较难区域,典型表现为在未见队列中漏检小的卫星病灶。在合成研究中,3种子成员间的分歧比单模型置信度更敏感地反映扫描差异。但随着扰动严重度增加,其体素级错误定位能力减弱。本工作提供严谨、诚实的可靠性对比,不主张任何单一不确定性方法占优。
原文摘要 · Abstract (English)
Deep networks segment brain tumours accurately in-distribution, but can fail silently when the input differs from their training data. That risk is central to clinical deployment and is the premise of the BraTS-GoAT generalizability task. We ask not only how well a model segments, but whether its uncertainty knows when it is wrong. On BraTS-GoAT (Task 3) we train a 5-fold cross-validated nnU-Net baseline (one held-out prediction per case) and a 3-seed deep ensemble. Both are evaluated for calibration and error detection on a per-region relevant mask, aggregated per case. In-distribution the 3-seed ensemble improves modestly over the already strong single model on the same held-out split, with the clearest gain in calibration. The separation appears under shift. In a controlled robustness study using graded synthetic corruptions as a proxy for acquisition shift, the single model's confidence stays flat while its accuracy and calibration degrade. Inter-member disagreement instead rises steeply, about a quarter to a third above the clean condition, several times the single model's response. On the official validation leaderboard the 5-fold ensemble of those folds attains whole-tumour Dice 0.87. The generalization gap is concentrated on the harder regions, with a characteristic failure of missing small, satellite lesions on unseen cohorts. In the synthetic study, disagreement among the 3-seed members is a more sensitive case-level indicator of acquisition shift than single-model confidence. Its per-voxel error localisation weakens as severity grows. The contribution is a rigorous, honest reliability comparison rather than a claim that any one uncertainty method dominates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。