arXiv:2410.05905cs.CV2024-10被引 11

用提示驱动方法提升医学图像分割通用性,参数少效果优。

MedUniSeg: 2D and 3D Medical Image Segmentation via a Prompt-driven Universal Model

  • 通过模态与任务提示生成先验,系统化引入编码器首尾。
  • 多任务平均Dice提升1.2%,参数量不足基线的十分之一。
  • 适合需要跨模态、跨任务分割的医疗影像研究者使用。

通用分割模型在处理多样化任务时具有巨大潜力,尤其依赖离散标注。随着任务和模态范围扩大,如何生成并合理部署任务与模态相关的先验成为关键。现有模型常忽略先验间的关联,且其位置与频率配置尚未深入探索。本文提出MedUniSeg,一种用于2D/3D多任务医学图像分割的提示驱动通用模型。该模型采用多个模态特定提示与一个通用任务提示,精确表征模态与任务特征。为生成相关先验,设计了模态图(MMap)与融合选择(FUSE)模块,将提示转化为对应先验,并系统地置于编码过程起始与末尾。在包含17个子数据集的多模态上游数据集上评估显示,MedUniSeg在17项任务上的平均Dice分数较nnUNet基线提升1.2%,参数量低于其1/10。对初始联合训练中表现不佳的任务,冻结MedUniSeg并引入新模块重学,得到增强版MedUniSeg*,在所有任务上均优于原版。此外,MedUniSeg在六个下游任务上超越先进自监督与监督预训练模型,展现出高泛化能力。

原文摘要 · Abstract (English)

Universal segmentation models offer significant potential in addressing a wide range of tasks by effectively leveraging discrete annotations. As the scope of tasks and modalities expands, it becomes increasingly important to generate and strategically position task- and modal-specific priors within the universal model. However, existing universal models often overlook the correlations between different priors, and the optimal placement and frequency of these priors remain underexplored. In this paper, we introduce MedUniSeg, a prompt-driven universal segmentation model designed for 2D and 3D multi-task segmentation across diverse modalities and domains. MedUniSeg employs multiple modal-specific prompts alongside a universal task prompt to accurately characterize the modalities and tasks. To generate the related priors, we propose the modal map (MMap) and the fusion and selection (FUSE) modules, which transform modal and task prompts into corresponding priors. These modal and task priors are systematically introduced at the start and end of the encoding process. We evaluate MedUniSeg on a comprehensive multi-modal upstream dataset consisting of 17 sub-datasets. The results demonstrate that MedUniSeg achieves superior multi-task segmentation performance, attaining a 1.2% improvement in the mean Dice score across the 17 upstream tasks compared to nnUNet baselines, while using less than 1/10 of the parameters. For tasks that underperform during the initial multi-task joint training, we freeze MedUniSeg and introduce new modules to re-learn these tasks. This approach yields an enhanced version, MedUniSeg*, which consistently outperforms MedUniSeg across all tasks. Moreover, MedUniSeg surpasses advanced self-supervised and supervised pre-trained models on six downstream tasks, establishing itself as a high-quality, highly generalizable pre-trained segmentation model.

医学图像分割提示学习通用模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。