用SAM思想做3D医学图像少样本分割,效率更高更准。
SAM-Guided Robust Representation Learning for One-Shot 3D Medical Image Segmentation
- 用双阶段知识蒸馏让轻量模型学懂SAM的通用特征
- 互指数移动平均+伪标签协同训练,提升分割精度
- 自动生成提示词,适合医疗数据少、算力有限场景
少样本医学图像分割对降低专家标注负担至关重要。尽管段落一切模型(SAM)在医学分割中展现出强大泛化能力,但其依赖人工交互和高计算成本限制了在少样本3D医学图像分割中的直接应用。为此,我们提出一种新型的SAM引导鲁棒表征学习框架RRL-MedSAM,利用SAM编码器的强大泛化能力来学习更优特征表示。设计双阶段知识蒸馏(DSKD)策略,从基础模型中蒸馏自然图像与医学图像间的通用知识,训练轻量编码器;再通过互指数移动平均(mutual-EMA)更新通用轻量编码器与医学专用编码器权重,其中注册网络生成的伪标签实现两者之间的相互监督。此外,引入自动提示(AP)分割解码器,将通用轻量模型生成的掩码作为提示,辅助医学专用模型提升最终分割性能。在三个公开数据集OASIS、CT-lung上进行的大量实验表明,所提RRL-MedSAM在分割与配准任务上均优于现有最先进方法。特别地,轻量编码器参数量仅为SAM-Base编码器的3%。
原文摘要 · Abstract (English)
One-shot medical image segmentation (MIS) is crucial for medical analysis due to the burden of medical experts on manual annotation. The recent emergence of the segment anything model (SAM) has demonstrated remarkable adaptation in MIS but cannot be directly applied to one-shot medical image segmentation (MIS) due to its reliance on labor-intensive user interactions and the high computational cost. To cope with these limitations, we propose a novel SAM-guided robust representation learning framework, named RRL-MedSAM, to adapt SAM to one-shot 3D MIS, which exploits the strong generalization capabilities of the SAM encoder to learn better feature representation. We devise a dual-stage knowledge distillation (DSKD) strategy to distill general knowledge between natural and medical images from the foundation model to train a lightweight encoder, and then adopt a mutual exponential moving average (mutual-EMA) to update the weights of the general lightweight encoder and medical-specific encoder. Specifically, pseudo labels from the registration network are used to perform mutual supervision for such two encoders. Moreover, we introduce an auto-prompting (AP) segmentation decoder which adopts the mask generated from the general lightweight model as a prompt to assist the medical-specific model in boosting the final segmentation performance. Extensive experiments conducted on three public datasets, i.e., OASIS, CT-lung demonstrate that the proposed RRL-MedSAM outperforms state-of-the-art one-shot MIS methods for both segmentation and registration tasks. Especially, our lightweight encoder uses only 3\% of the parameters compared to the encoder of SAM-Base.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。