将视频模型SAM2改进用于3D医学图像分割,提升边界精度与跨切片连续性建模。
SAM2-3dMed: Empowering SAM2 for 3D Medical Image Segmentation
- 通过自监督预测切片相对位置,建模3D图像中双向切片依赖关系。
- 引入边界检测模块,在肺、脾、胰腺数据集上平均Dice达0.92以上。
- 适合需要高精度解剖边界分割的临床研究与医学影像分析任务。
准确分割3D医学图像对疾病评估和治疗规划至关重要。尽管分割一切模型2(SAM2)在利用时序信息进行视频目标分割方面表现卓越,但其直接应用于3D医学图像存在两大根本性领域差异:1)切片间具有双向解剖连续性,与视频中的单向时序流截然不同;2)精确的边界划分对形态学分析至关重要,但在视频任务中常被忽略。为此,我们提出SAM2-3dMed,一种针对3D医学影像的SAM2适配框架。本框架引入两项关键创新:1)切片相对位置预测(SRPP)模块,通过自监督方式显式建模切片间的双向依赖关系;2)边界检测(BD)模块,增强关键器官与组织边界的分割精度。在三个多样化医学数据集(医学分割十项全能赛中的肺、脾、胰腺)上的大量实验表明,SAM2-3dMed显著优于现有先进方法,在分割重叠率与边界精度上均取得优异表现。该方法不仅提升了3D医学图像分割性能,还为将面向视频的通用模型迁移到空间体数据提供了通用范式。
原文摘要 · Abstract (English)
Accurate segmentation of 3D medical images is critical for clinical applications like disease assessment and treatment planning. While the Segment Anything Model 2 (SAM2) has shown remarkable success in video object segmentation by leveraging temporal cues, its direct application to 3D medical images faces two fundamental domain gaps: 1) the bidirectional anatomical continuity between slices contrasts sharply with the unidirectional temporal flow in videos, and 2) precise boundary delineation, crucial for morphological analysis, is often underexplored in video tasks. To bridge these gaps, we propose SAM2-3dMed, an adaptation of SAM2 for 3D medical imaging. Our framework introduces two key innovations: 1) a Slice Relative Position Prediction (SRPP) module explicitly models bidirectional inter-slice dependencies by guiding SAM2 to predict the relative positions of different slices in a self-supervised manner; 2) a Boundary Detection (BD) module enhances segmentation accuracy along critical organ and tissue boundaries. Extensive experiments on three diverse medical datasets (the Lung, Spleen, and Pancreas in the Medical Segmentation Decathlon (MSD) dataset) demonstrate that SAM2-3dMed significantly outperforms state-of-the-art methods, achieving superior performance in segmentation overlap and boundary precision. Our approach not only advances 3D medical image segmentation performance but also offers a general paradigm for adapting video-centric foundation models to spatial volumetric data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。