arXiv:2505.12203eess.IVcs.CV2025-05被引 13

融合卷积与注意力机制,提升低剂量CT图像去噪效果

CTLformer: A Hybrid Denoising Model Combining Convolutional Layers and Self-Attention for Enhanced CT Image Reconstruction

  • 结合卷积与Transformer,用多尺度注意力捕捉细节和全局结构
  • 动态调整注意力分布,高噪声区重点降噪,低噪声区保留细节
  • 在权威数据集上性能领先,适合临床复杂噪声场景

低剂量CT(LDCT)图像常伴随显著噪声,影响图像质量与诊断准确性。为应对LDCT去噪中多尺度特征融合与多样噪声分布的挑战,本文提出一种新型混合模型CTLformer,融合卷积结构与Transformer架构。核心创新包括:多尺度注意力机制,通过Token2Token机制与自注意力交互模块,有效捕获不同尺度下的精细细节与全局结构,增强有效特征并抑制噪声;动态注意力控制机制,根据输入图像噪声特性自适应调整注意力分布,在高噪声区域聚焦降噪,同时保护低噪声区域细节,提升鲁棒性与去噪性能。此外,CTLformer集成卷积层实现高效特征提取,并采用重叠推理策略减少边界伪影,进一步强化去噪能力。在2016年美国国立卫生研究院AAPM梅奥诊所低剂量CT挑战赛数据集上的实验结果表明,CTLformer在去噪性能与模型效率方面均显著优于现有方法,大幅提升了LDCT图像质量。所提模型不仅为LDCT去噪提供了高效解决方案,更在医学图像分析领域,尤其针对复杂噪声模式的临床应用中展现出广泛潜力。

原文摘要 · Abstract (English)

Low-dose CT (LDCT) images are often accompanied by significant noise, which negatively impacts image quality and subsequent diagnostic accuracy. To address the challenges of multi-scale feature fusion and diverse noise distribution patterns in LDCT denoising, this paper introduces an innovative model, CTLformer, which combines convolutional structures with transformer architecture. Two key innovations are proposed: a multi-scale attention mechanism and a dynamic attention control mechanism. The multi-scale attention mechanism, implemented through the Token2Token mechanism and self-attention interaction modules, effectively captures both fine details and global structures at different scales, enhancing relevant features and suppressing noise. The dynamic attention control mechanism adapts the attention distribution based on the noise characteristics of the input image, focusing on high-noise regions while preserving details in low-noise areas, thereby enhancing robustness and improving denoising performance. Furthermore, CTLformer integrates convolutional layers for efficient feature extraction and uses overlapping inference to mitigate boundary artifacts, further strengthening its denoising capability. Experimental results on the 2016 National Institutes of Health AAPM Mayo Clinic LDCT Challenge dataset demonstrate that CTLformer significantly outperforms existing methods in both denoising performance and model efficiency, greatly improving the quality of LDCT images. The proposed CTLformer not only provides an efficient solution for LDCT denoising but also shows broad potential in medical image analysis, especially for clinical applications dealing with complex noise patterns.

CT图像重建去噪Transformer医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。