通过分层隐变量建模,提升多中心前列腺病灶分割的泛化能力。
Deep EM with Hierarchical Latent Label Modelling for Multi-Site Prostate Lesion Segmentation
- 用分层期望最大化框架,将标注视为噪声观测,推断潜在清洁掩码
- 跨中心分割DSC提升至27.91%~39.69%,显著优于现有方法(p<0.039)
- 可生成可解释的各中心标注质量估计,支持后续变异分析
标注差异是前列腺病灶分割的主要挑战。在多中心数据集中,标注常反映中心特有的勾画协议,导致分割网络过拟合于局部风格,推理时对未见中心泛化能力差。本文将每个观测标注视为潜在‘干净’病灶掩码的噪声观测,提出分层期望最大化(HierEM)框架,交替执行:(1) 推断体素级潜在掩码后验分布;(2) 使用该后验作为软目标训练CNN,并在分层先验下估计各中心的敏感性与特异性。该先验将标注质量分解为全局均值及中心和病例级偏差,通过惩罚仅由中心偏差贡献的似然项,减少中心偏差影响。在三个队列上的实验表明,所提方法在跨中心泛化性能上优于现有先进方法。合并数据集评估中,各中心平均DSC为29.50%~39.69%;留一中心外验证中为27.91%~32.67%,统计显著优于对比方法(p<0.039)。方法还生成可解释的各中心潜标注质量估计(敏感性α为31.5%~47.3%,特异性β≈0.99),支持对跨中心标注变异的后处理分析。结果表明,显式建模中心依赖标注可提升跨中心泛化能力。
原文摘要 · Abstract (English)
Label variability is a major challenge for prostate lesion segmentation. In multi-site datasets, annotations often reflect centre-specific contouring protocols, causing segmentation networks to overfit to local styles and generalise poorly to unseen sites in inference. We treat each observed annotation as a noisy observation of an underlying latent 'clean' lesion mask, and propose a hierarchical expectation-maximisation (HierEM) framework that alternates between: (1) inferring a voxel-wise posterior distribution over the latent mask, and (2) training a CNN using this posterior as a soft target and estimate site-specific sensitivity and specificity under a hierarchical prior. This hierarchical prior decomposes label-quality into a global mean with site- and case-level deviations, reducing site-specific bias by penalising the likelihood term contributed only by site deviations. Experiments on three cohorts demonstrate that the proposed hierarchical EM framework enhances cross-site generalisation compared to state-of-the-art methods. For pooled-dataset evaluation, the per-site mean DSC ranges from 29.50% to 39.69%; for leave-one-site-out generalisation, it ranges from 27.91% to 32.67%, yielding statistically significant improvements over comparison methods (p<0.039). The method also produces interpretable per-site latent label-quality estimates (sensitivity alpha ranges from 31.5% to 47.3% at specificity beta approximates 0.99), supporting post-hoc analyses of cross-site annotation variability. These results indicate that explicitly modelling site-dependent annotation can improve cross-site generalisation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。