arXiv:2412.05605cs.CV2024-12被引 2

将SAM扩展到3D医学影像分割,提升器官与病灶识别精度

RefSAM3D: Adapting SAM with Cross-modal Reference for 3D Medical Image Segmentation

  • 用3D图像适配器和跨模态文本提示增强模型对体积数据的建模能力
  • 在多个医疗数据集上优于现有方法,复杂解剖结构分割更准确
  • 适合需要高精度3D影像分析的医学研究人员与临床医生

原始的通用图像分割模型SAM基于2D视觉变换器(ViT),擅长处理自然图像中的全局模式,但在CT、MRI等3D医学影像任务中表现不佳。这类任务需在体积分空间中捕捉精确的空间信息,以实现器官分割和肿瘤量化。为此,我们提出RefSAM3D,通过引入3D图像适配器和跨模态参考提示生成机制,将SAM拓展至3D医学影像。该方法改造了视觉编码器以处理3D输入,并优化掩码解码器实现直接3D掩码生成。同时融合文本提示以提升复杂解剖场景下的分割精度与一致性。通过层次化注意力机制,有效整合多尺度信息。在多个医学影像数据集上的广泛评估表明,RefSAM3D显著优于当前先进方法。本工作推动了SAM在复杂解剖结构精准分割中的应用。

原文摘要 · Abstract (English)

The Segment Anything Model (SAM), originally built on a 2D Vision Transformer (ViT), excels at capturing global patterns in 2D natural images but struggles with 3D medical imaging modalities like CT and MRI. These modalities require capturing spatial information in volumetric space for tasks such as organ segmentation and tumor quantification. To address this challenge, we introduce RefSAM3D, which adapts SAM for 3D medical imaging by incorporating a 3D image adapter and cross-modal reference prompt generation. Our approach modifies the visual encoder to handle 3D inputs and enhances the mask decoder for direct 3D mask generation. We also integrate textual prompts to improve segmentation accuracy and consistency in complex anatomical scenarios. By employing a hierarchical attention mechanism, our model effectively captures and integrates information across different scales. Extensive evaluations on multiple medical imaging datasets demonstrate the superior performance of RefSAM3D over state-of-the-art methods. Our contributions advance the application of SAM in accurately segmenting complex anatomical structures in medical imaging.

3D分割医学影像跨模态SAM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。