医学影像分割中,共识方法常不如简单投票,需谨慎选择。
When Does Consensus Beat Voting? A Critical Analysis of Statistical Label Fusion in Medical Image Segmentation
- 从生成模型出发推导共识算法,揭示其理论局限
- 实验证明STAPLE在多数情况下等效于阈值投票,优化率仅5%
- 建议用深度共识+置信区间提升可靠性,避免盲目使用传统方法
本文对共识分割进行了严格、自包含的分析。从生成模型、EM算法、Van Leemput边缘化分析、可识别性条件、空间STAPLE及深度变分公式出发,通过受控实验验证了每项理论预测。核心发现令人警醒:在常见条件下,STAPLE退化为阈值多数投票,其期望最大化(EM)优化效率仅5%,且在类别不平衡下严重失效。这些并非极端情况,而是医学影像中的典型场景。多数投票——简单、非参数、鲁棒——是出人意料的强大基线,领域可能过早抛弃它以追求更“复杂”的方法。同时,深度共识模型表明,当图像与标签联合使用时,共识问题并非固有困难;置信区间方法也证明了形式化不确定性保证可实现且实用。我们希望该工作促使从业者批判性评估共识方法,而非默认使用STAPLE,为更严谨的方法提供数学与实证基础。
原文摘要 · Abstract (English)
This paper provides a rigorous, self-contained investigation of consensus segmentation. We derive the mathematical foundations from first principles -- the generative model, EM algorithm, Van Leemput's marginalization analysis, identifiability conditions, Spatial STAPLE, and deep variational formulations -- and validate each theoretical prediction through controlled experiments. The central finding is sobering: under common conditions, STAPLE reduces to thresholded majority voting, suffers 95% EM suboptimality, and collapses under class imbalance. These are not edge cases but typical scenarios in medical imaging. Majority voting -- simple, non-parametric, and robust -- is a surprisingly strong baseline that the field has perhaps too hastily dismissed in favor of more "sophisticated" methods. At the same time, the deep consensus model demonstrates that the consensus problem is not inherently difficult -- it becomes tractable when the image is used alongside the labels. And conformal prediction shows that formal uncertainty guarantees are achievable and practical. We hope this work encourages practitioners to critically evaluate their consensus methods rather than applying STAPLE by default, and provides the mathematical and empirical foundation for more principled approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。