arXiv:2409.17675cs.CV2024-09中稿 · MICCAI 2024被引 41

用Mamba提升3D医学图像分割效率与精度

EM-Net: Efficient Channel and Frequency Learning with Mamba for 3D Medical Image Segmentation

  • 融合通道选择与频率域分析,增强跨尺度特征学习
  • 参数量减半,训练速度提升2倍,精度优于现有模型
  • 适合需要高效高精度分割的医学影像研究者

卷积神经网络在3D医学图像分割中占据主导地位,但受限于较小的感受野。变换器模型虽能捕捉全局关系,但在高分辨率下计算成本高昂。近期出现的状态空间模型Mamba在序列建模方面表现优异。受此启发,我们提出一种基于Mamba的新型3D医学图像分割模型EM-Net。该模型通过整合与选择通道,高效捕获区域间的注意力交互;同时利用频率域有效协调不同尺度特征的学习,显著提升训练速度。在两个具有挑战性的多器官数据集上,与当前最优(SOTA)方法相比,本方法在保持更高分割精度的同时,参数量仅需SOTA模型的一半,训练速度提升2倍。

原文摘要 · Abstract (English)

Convolutional neural networks have primarily led 3D medical image segmentation but may be limited by small receptive fields. Transformer models excel in capturing global relationships through self-attention but are challenged by high computational costs at high resolutions. Recently, Mamba, a state space model, has emerged as an effective approach for sequential modeling. Inspired by its success, we introduce a novel Mamba-based 3D medical image segmentation model called EM-Net. It not only efficiently captures attentive interaction between regions by integrating and selecting channels, but also effectively utilizes frequency domain to harmonize the learning of features across varying scales, while accelerating training speed. Comprehensive experiments on two challenging multi-organ datasets with other state-of-the-art (SOTA) algorithms show that our method exhibits better segmentation accuracy while requiring nearly half the parameter size of SOTA models and 2x faster training speed.

3D分割Mamba医学图像高效模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。