用全局轴向注意力提升3D医学图像分割精度,尤其擅长小器官和模糊边界。
GASA-UNet: Global Axial Self-Attention U-Net for 3D Medical Image Segmentation
- 设计新型全局轴向自注意力块,跨平面建模3D医学图像
- 在BTCV、AMOS、KiTS23数据集上显著提升Dice与NSD得分
- 适合需要高精度分割小器官或病变的临床研究者
准确分割多器官及病理组织在医学影像中至关重要但极具挑战,尤其面对细微分类与模糊边界。为此,我们提出GASA-UNet,一种改进型U-Net模型,引入新型全局轴向自注意力(GASA)模块。该模块将图像作为3D实体处理,每个2D切片代表不同解剖截面,利用提取的1D片段上的多头自注意力(MHSA)机制,在各切片间建立关联。同时在注意力框架中加入位置编码(PE),增强体素特征的空间上下文信息,提升组织分类与器官边缘识别能力。模型在三个基准数据集BTCV、AMOS和KiTS23上表现出色,对小型解剖结构的分割性能显著提升,验证了其在Dice分数和归一化表面Dice(NSD)指标上的优越性。
原文摘要 · Abstract (English)
Accurate segmentation of multiple organs and the differentiation of pathological tissues in medical imaging are crucial but challenging, especially for nuanced classifications and ambiguous organ boundaries. To tackle these challenges, we introduce GASA-UNet, a refined U-Net-like model featuring a novel Global Axial Self-Attention (GASA) block. This block processes image data as a 3D entity, with each 2D plane representing a different anatomical cross-section. Voxel features are defined within this spatial context, and a Multi-Head Self-Attention (MHSA) mechanism is utilized on extracted 1D patches to facilitate connections across these planes. Positional embeddings (PE) are incorporated into our attention framework, enriching voxel features with spatial context and enhancing tissue classification and organ edge delineation. Our model has demonstrated promising improvements in segmentation performance, particularly for smaller anatomical structures, as evidenced by enhanced Dice scores and Normalized Surface Dice (NSD) on three benchmark datasets, i.e., BTCV, AMOS, and KiTS23.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。