提出S3TU-Net模型,提升肺结节在CT图像中的分割精度。
S3TU-Net: Structured Convolution and Superpixel Transformer for Lung Nodule Segmentation
- 融合超像素视觉变压器与结构化卷积,捕捉多尺度特征。
- 在LIDC-IDRI数据集上达89.04% Dice系数,比现有方法高4.52%。
- 适合医学影像分割研究者,尤其关注肺结节检测的临床应用。
肺腺癌结节在CT图像中形态不规则、边界模糊,影响分期诊断,精准分割对临床提取病灶信息至关重要。本文提出S3TU-Net模型,结合多维空间连接器与基于超像素的视觉变压器。该模型采用多视角CNN-Transformer混合架构,融入超像素算法、结构化加权与空间移位技术,实现优异分割性能。通过结构化卷积模块(DWF-Conv/D2BR-Conv)提取多尺度局部特征并缓解过拟合。为增强多尺度特征融合,引入S2-MLP Link,在跳跃连接处集成空间移位与注意力机制。此外,基于残差的超像素视觉变压器(RM-SViT)利用稀疏相关学习与多分支注意力,有效融合全局与局部特征,捕获长程依赖,残差连接提升稳定性与计算效率。在LIDC-IDRI数据集上的实验表明,S3TU-Net达到89.04%的Dice系数、90.73%的精确率和90.70%的IoU。相比近期方法,其Dice系数提升4.52%,敏感性提高3.16%,其他指标约增2%。此外,在EPDB私有数据集上验证了泛化能力,取得86.40%的Dice系数。
原文摘要 · Abstract (English)
The irregular and challenging characteristics of lung adenocarcinoma nodules in computed tomography (CT) images complicate staging diagnosis, making accurate segmentation critical for clinicians to extract detailed lesion information. In this study, we propose a segmentation model, S3TU-Net, which integrates multi-dimensional spatial connectors and a superpixel-based visual transformer. S3TU-Net is built on a multi-view CNN-Transformer hybrid architecture, incorporating superpixel algorithms, structured weighting, and spatial shifting techniques to achieve superior segmentation performance. The model leverages structured convolution blocks (DWF-Conv/D2BR-Conv) to extract multi-scale local features while mitigating overfitting. To enhance multi-scale feature fusion, we introduce the S2-MLP Link, integrating spatial shifting and attention mechanisms at the skip connections. Additionally, the residual-based superpixel visual transformer (RM-SViT) effectively merges global and local features by employing sparse correlation learning and multi-branch attention to capture long-range dependencies, with residual connections enhancing stability and computational efficiency. Experimental results on the LIDC-IDRI dataset demonstrate that S3TU-Net achieves a DSC, precision, and IoU of 89.04%, 90.73%, and 90.70%, respectively. Compared to recent methods, S3TU-Net improves DSC by 4.52% and sensitivity by 3.16%, with other metrics showing an approximate 2% increase. In addition to comparison and ablation studies, we validated the generalization ability of our model on the EPDB private dataset, achieving a DSC of 86.40%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。