为医学影像交互分割设计临床真实评估方法,提升算法可靠性。
A methodology for clinically driven interactive segmentation evaluation
- 基于临床场景构建标准化评估流程,确保任务与指标真实可信。
- 发现用户交互信息丢失会显著降低模型鲁棒性,自适应缩放能加速收敛。
- 适合医学影像分割研究者、临床算法开发者使用,尤其关注真实应用效果。
交互式分割是实现体积医学图像分割鲁棒、通用算法的有前景策略。然而,不一致且脱离临床实际的评估方式阻碍了公平比较,并扭曲了真实表现。本文提出一种基于临床实践的评估任务与指标定义方法,并构建了标准化评估流水线软件框架。我们在异构复杂任务中评估了前沿算法,发现:(i) 处理用户交互时最小化信息损失对模型鲁棒性至关重要;(ii) 自适应缩放机制可提升鲁棒性并加快收敛速度;(iii) 若验证阶段的提示行为或预算与训练阶段不同,性能会下降;(iv) 2D 方法在层状图像和粗略目标上表现良好,而3D上下文对大或不规则形状的目标更有效;(v) 非医学领域模型(如SAM2)在对比度差或形状复杂时性能显著下降。
原文摘要 · Abstract (English)
Interactive segmentation is a promising strategy for building robust, generalisable algorithms for volumetric medical image segmentation. However, inconsistent and clinically unrealistic evaluation hinders fair comparison and misrepresents real-world performance. We propose a clinically grounded methodology for defining evaluation tasks and metrics, and built a software framework for constructing standardised evaluation pipelines. We evaluate state-of-the-art algorithms across heterogeneous and complex tasks and observe that (i) minimising information loss when processing user interactions is critical for model robustness, (ii) adaptive-zooming mechanisms boost robustness and speed convergence, (iii) performance drops if validation prompting behaviour/budgets differ from training, (iv) 2D methods perform well with slab-like images and coarse targets, but 3D context helps with large or irregularly shaped targets, (v) performance of non-medical-domain models (e.g. SAM2) degrades with poor contrast and complex shapes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。