提出3D多模态脑肿瘤分割新模型,提升小病灶识别精度。
Multi-Modal Brain Tumor Segmentation via 3D Multi-Scale Self-attention and Cross-attention
- 通过多尺度分块与聚合,在自注意力中同时捕捉多尺度特征和长程依赖。
- 在三个公开数据集上,平均Dice分数优于现有最优方法。
- 适合需要精准分割微小肿瘤的临床医学影像分析场景。
由于基于CNN和Transformer的模型在各类计算机视觉任务中表现优异,近期研究探索了混合架构在3D多模态医学图像分割中的应用。引入Transformer使混合模型具备建模3D医学图像长距离依赖的能力。然而,这些模型通常在每个自注意力层中采用固定的3D体素感受野,忽略了多尺度病灶特征。为此,我们提出一种基于编码器-解码器结构的CNN-Transformer混合3D医学图像分割模型TMA-TransBTS。该模型通过自注意力层中3D标记的多尺度划分与聚合,实现多尺度3D特征的同时提取与长距离依赖建模。此外,TMA-TransBTS设计了一种3D多尺度交叉注意力模块,利用交叉注意力机制与3D标记的多尺度聚合,在编码器与解码器间建立联系,以提取丰富的体积表征。在三个公开3D医学分割数据集上的大量实验表明,TMA-TransBTS在3D多模态脑肿瘤分割任务中,平均分割性能优于先前最先进的基于CNN的3D方法及混合3D方法。
原文摘要 · Abstract (English)
Due to the success of CNN-based and Transformer-based models in various computer vision tasks, recent works study the applicability of CNN-Transformer hybrid architecture models in 3D multi-modality medical segmentation tasks. Introducing Transformer brings long-range dependent information modeling ability in 3D medical images to hybrid models via the self-attention mechanism. However, these models usually employ fixed receptive fields of 3D volumetric features within each self-attention layer, ignoring the multi-scale volumetric lesion features. To address this issue, we propose a CNN-Transformer hybrid 3D medical image segmentation model, named TMA-TransBTS, based on an encoder-decoder structure. TMA-TransBTS realizes simultaneous extraction of multi-scale 3D features and modeling of long-distance dependencies by multi-scale division and aggregation of 3D tokens in a self-attention layer. Furthermore, TMA-TransBTS proposes a 3D multi-scale cross-attention module to establish a link between the encoder and the decoder for extracting rich volume representations by exploiting the mutual attention mechanism of cross-attention and multi-scale aggregation of 3D tokens. Extensive experimental results on three public 3D medical segmentation datasets show that TMA-TransBTS achieves higher averaged segmentation results than previous state-of-the-art CNN-based 3D methods and CNN-Transform hybrid 3D methods for the segmentation of 3D multi-modality brain tumors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。