用状态空间模型提升低剂量CT图像去噪效果,兼顾远距离与局部细节。
DenoMamba: A fused state-space model for low-dose CT denoising
- 融合空间与通道状态空间模块,高效捕捉图像长程与短程上下文。
- 在25%和10%剂量下,平均提升1.4dB PSNR、1.1% SSIM、1.6% RMSE。
- 适合医学影像去噪场景,尤其适用于高分辨率低剂量CT重建。
低剂量计算机断层扫描(LDCT)可降低辐射风险,但需依赖先进去噪算法维持图像诊断质量。当前主流方法基于神经网络学习数据驱动的图像先验,以分离剂量降低引起的噪声与组织信号。然而,这些先验的有效性取决于模型对复杂上下文特征的捕捉能力。早期卷积神经网络(CNN)擅长捕捉短距离空间上下文,但受限于感受野,对长距离交互不敏感;尽管基于自注意力机制的变压器能增强长程感知,却因模型复杂度高导致效率下降,尤其在高分辨率CT图像上表现不佳。为此,本文提出DenoMamba,一种基于状态空间模型(SSM)的新型去噪方法,可高效捕获医学图像中的短程与长程上下文。其采用编码器-解码器结构,结合空间SSM模块编码空间上下文,以及配备二级门控卷积网络的通道SSM模块,用于编码各阶段的通道特征。两个模块的特征图通过卷积融合模块(CFM)与低层输入特征融合。在25%与10%剂量减少的LDCT数据集上的全面实验表明,DenoMamba显著优于现有最优去噪器,图像恢复质量平均提升1.4dB PSNR、1.1% SSIM、1.6% RMSE。
原文摘要 · Abstract (English)
Low-dose computed tomography (LDCT) lower potential risks linked to radiation exposure while relying on advanced denoising algorithms to maintain diagnostic quality in reconstructed images. The reigning paradigm in LDCT denoising is based on neural network models that learn data-driven image priors to separate noise evoked by dose reduction from underlying tissue signals. Naturally, the fidelity of these priors depend on the model's ability to capture the broad range of contextual features evident in CT images. Earlier convolutional neural networks (CNN) are highly adept at efficiently capturing short-range spatial context, but their limited receptive fields reduce sensitivity to interactions over longer distances. Although transformers based on self-attention mechanisms have recently been posed to increase sensitivity to long-range context, they can suffer from suboptimal performance and efficiency due to elevated model complexity, particularly for high-resolution CT images. For high-quality restoration of LDCT images, here we introduce DenoMamba, a novel denoising method based on state-space modeling (SSM), that efficiently captures short- and long-range context in medical images. Following an hourglass architecture with encoder-decoder stages, DenoMamba employs a spatial SSM module to encode spatial context and a novel channel SSM module equipped with a secondary gated convolution network to encode latent features of channel context at each stage. Feature maps from the two modules are then consolidated with low-level input features via a convolution fusion module (CFM). Comprehensive experiments on LDCT datasets with 25\% and 10\% dose reduction demonstrate that DenoMamba outperforms state-of-the-art denoisers with average improvements of 1.4dB PSNR, 1.1% SSIM, and 1.6% RMSE in recovered image quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。