arXiv:2502.13524cs.CVcs.AI2025-02

轻量级3D医学图像分割模型,速度超90帧/秒且精度领先。

MobileViM: A Light-weight and Dimension-independent Vision Mamba for 3D Medical Image Analysis

  • 设计维度无关机制与双向遍历结构,提升3D医学图像处理效率。
  • 单卡实测速度达90+帧/秒,比现有模型快24帧以上。
  • 适用于高精度医疗影像分析,尤其适合实时诊断场景。

三维(3D)医学图像的高效评估对医疗诊断与治疗至关重要。近年来,深度学习与计算机视觉在医学图像分析中广泛应用。传统方法如卷积神经网络(CNN)和视觉变换器(ViTs)面临显著计算挑战,亟需架构革新。近期出现的Mamba模型在处理一维数据时具有低计算开销的优势,但其在3D医学图像分析中的潜力尚未充分挖掘,且随着维度增加可能遭遇计算瓶颈。本文提出MobileViM,一种面向3D医学图像高效分割的轻量化架构。MobileViM引入新型维度无关机制与双方向遍历策略,结合基于视觉Mamba的框架,并采用跨尺度桥接技术以提升多模态医学影像下的效率与精度。实验表明,MobileViM在单张NVIDIA RTX 4090 GPU上实现超过90帧/秒的分割速度,较当前最优模型快24帧以上。在PENGWIN、BraTS2024、ATLAS和Toothfairy2数据集上的Dice相似系数分别达到92.72%、86.69%、80.46%和77.43%,显著优于现有模型。

原文摘要 · Abstract (English)

Efficient evaluation of three-dimensional (3D) medical images is crucial for diagnostic and therapeutic practices in healthcare. Recent years have seen a substantial uptake in applying deep learning and computer vision to analyse and interpret medical images. Traditional approaches, such as convolutional neural networks (CNNs) and vision transformers (ViTs), face significant computational challenges, prompting the need for architectural advancements. Recent efforts have led to the introduction of novel architectures like the ``Mamba'' model as alternative solutions to traditional CNNs or ViTs. The Mamba model excels in the linear processing of one-dimensional data with low computational demands. However, Mamba's potential for 3D medical image analysis remains underexplored and could face significant computational challenges as the dimension increases. This manuscript presents MobileViM, a streamlined architecture for efficient segmentation of 3D medical images. In the MobileViM network, we invent a new dimension-independent mechanism and a dual-direction traversing approach to incorporate with a vision-Mamba-based framework. MobileViM also features a cross-scale bridging technique to improve efficiency and accuracy across various medical imaging modalities. With these enhancements, MobileViM achieves segmentation speeds exceeding 90 frames per second (FPS) on a single graphics processing unit (i.e., NVIDIA RTX 4090). This performance is over 24 FPS faster than the state-of-the-art deep learning models for processing 3D images with the same computational resources. In addition, experimental evaluations demonstrate that MobileViM delivers superior performance, with Dice similarity scores reaching 92.72%, 86.69%, 80.46%, and 77.43% for PENGWIN, BraTS2024, ATLAS, and Toothfairy2 datasets, respectively, which significantly surpasses existing models.

3D医学图像轻量级模型Mamba分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。