用Transformer模型提升超声弹性成像的清晰度与分割精度
SW-ViT: A Spatio-Temporal Vision Transformer Network with Post Denoiser for Sequential Multi-Push Ultrasound Shear Wave Elastography
- 分两阶段设计:先重建弹性图,再通过去噪网络精修并生成分割图
- 模拟数据下结构相似性达0.995,实际模型测试分割准确率0.738
- 适合医学影像处理、超声诊断领域研究者参考
目的:超声剪切波弹性成像(SWE)通过映射组织硬度评估软组织病变,与恶性程度相关。传统方法存在噪声敏感、训练数据少、无法同时生成分割掩码等问题。本文提出SW-ViT,一种两阶段深度学习框架,结合基于CNN-时空视觉变换器的重建网络和高效的Transformer后置去噪网络。第一阶段采用3D ResNet编码器搭配多分辨率时空变换块,捕捉时空特征,经压缩-激励注意力解码器重建二维硬度图;为缓解数据不足,采用基于图像块的局部学习策略。第二阶段使用共享编码器与双解码器,分别处理包含物和背景区域,输出优化后的硬度图与分割掩码。混合损失函数融合区域、平滑性、融合与交并比(IoU)成分,提升重建与分割性能。在模拟数据上,本方法获得PSNR 32.68 dB、CNR 46.78 dB、SSIM 0.995;在体模数据上,分别为PSNR 21.11 dB、CNR 42.14 dB、SSIM 0.936;分割IoU达0.949(模拟)与0.738(体模),平均表面距离(ASSD)分别为0.184和1.011。结果表明,SW-ViT能从噪声干扰的SWE数据中生成鲁棒且高质量的弹性图,具有明确临床应用前景。
原文摘要 · Abstract (English)
Objective: Ultrasound Shear Wave Elastography (SWE) demonstrates great potential in assessing soft-tissue pathology by mapping tissue stiffness, which is linked to malignancy. Traditional SWE methods have shown promise in estimating tissue elasticity, yet their susceptibility to noise interference, reliance on limited training data, and inability to generate segmentation masks concurrently present notable challenges to accuracy and reliability. Approach: In this paper, we propose SW-ViT, a novel two-stage deep learning framework for SWE that integrates a CNN-Spatio-Temporal Vision Transformer-based reconstruction network with an efficient Transformer-based post-denoising network. The first stage uses a 3D ResNet encoder with multi-resolution spatio-temporal Transformer blocks that capture spatial and temporal features, followed by a squeeze-and-excitation attention decoder that reconstructs 2D stiffness maps. To address data limitations, a patch-based training strategy is adopted for localized learning and reconstruction. In the second stage, a denoising network with a shared encoder and dual decoders processes inclusion and background regions to produce a refined stiffness map and segmentation mask. A hybrid loss combining regional, smoothness, fusion, and Intersection over Union (IoU) components ensures improvements in both reconstruction and segmentation. Results: On simulated data, our method achieves PSNR of 32.68 dB, CNR of 46.78 dB, and SSIM of 0.995. On phantom data, results include PSNR of 21.11 dB, CNR of 42.14 dB, and SSIM of 0.936. Segmentation IoU values reach 0.949 (simulation) and 0.738 (phantom) with ASSD values being 0.184 and 1.011, respectively. Significance: SW-ViT delivers robust, high-quality elasticity map estimates from noisy SWE data and holds clear promise for clinical application.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。