arXiv:2412.10776eess.IVcs.AI2024-12AAAI被引 11

提升ViT在加速MRI重建中的表现,解决高频信息捕捉、冗余计算和多尺度建模难题。

Boosting ViT-based MRI Reconstruction from the Perspectives of Frequency Modulation, Spatial Purification, and Scale Diversification

  • 通过频率调制、空间净化和多尺度融合改进ViT结构
  • 在三个公开数据集上优于现有方法,且计算成本更低
  • 适合医学图像重建与高效Transformer设计的研究者

加速MRI重建因k空间大幅欠采样而成为典型的不适定逆问题。近年来,视觉变换器(ViTs)已成为主流方法,性能显著提升。但仍有三大问题未解决:(1) ViTs难以捕捉图像的高频成分,限制了局部纹理与边缘信息的恢复;(2) 以往方法在内容中对相关与无关令牌均计算多头自注意力,引入噪声并增加计算负担;(3) ViTs中简单的前馈网络无法有效建模对图像修复至关重要的多尺度信息。本文提出FPS-Former框架,从频率调制、空间净化和尺度多样化三方面应对上述挑战。具体而言,针对问题(1),设计频率调制注意力模块,通过拉普拉斯金字塔自适应重校准频率信息;针对问题(2),提出空间净化注意力模块,仅捕获紧密相关令牌间的交互,减少冗余特征表示;针对问题(3),构建基于混合尺度融合策略的高效前馈网络。在三个公开数据集上的全面实验表明,本方法在性能上超越现有最优方法,同时计算开销更低。

原文摘要 · Abstract (English)

The accelerated MRI reconstruction process presents a challenging ill-posed inverse problem due to the extensive under-sampling in k-space. Recently, Vision Transformers (ViTs) have become the mainstream for this task, demonstrating substantial performance improvements. However, there are still three significant issues remain unaddressed: (1) ViTs struggle to capture high-frequency components of images, limiting their ability to detect local textures and edge information, thereby impeding MRI restoration; (2) Previous methods calculate multi-head self-attention (MSA) among both related and unrelated tokens in content, introducing noise and significantly increasing computational burden; (3) The naive feed-forward network in ViTs cannot model the multi-scale information that is important for image restoration. In this paper, we propose FPS-Former, a powerful ViT-based framework, to address these issues from the perspectives of frequency modulation, spatial purification, and scale diversification. Specifically, for issue (1), we introduce a frequency modulation attention module to enhance the self-attention map by adaptively re-calibrating the frequency information in a Laplacian pyramid. For issue (2), we customize a spatial purification attention module to capture interactions among closely related tokens, thereby reducing redundant or irrelevant feature representations. For issue (3), we propose an efficient feed-forward network based on a hybrid-scale fusion strategy. Comprehensive experiments conducted on three public datasets show that our FPS-Former outperforms state-of-the-art methods while requiring lower computational costs.

MRI重建Vision Transformer频域建模多尺度融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。