用新型RWKV模型将3T核磁共振图像转为7T级高清图像,提升细节和对比度。
FS-RWKV: Leveraging Frequency Spatial-Aware RWKV for 3T-to-7T MRI Translation
- 基于RWKV架构,结合频率-空间感知模块增强全局上下文建模。
- 在两个数据集上优于主流方法,在解剖细节和感知质量上表现最佳。
- 适合医学影像合成、脑部疾病早期检测等临床场景使用。
超高场7T MRI能提供更高分辨率和组织对比度,有助于发现神经系统疾病的细微病变,但因设备成本高、技术要求严,临床普及受限。通过计算方法将易获取的3T图像转换为7T质量图像,是解决可及性问题的可行方案。现有CNN方法存在空间覆盖有限的问题,而Transformer模型计算开销过大。RWKV架构具备线性复杂度与强长程依赖捕捉能力,适合医疗图像合成。本文提出频率-空间感知的RWKV框架(FS-RWKV),包含两个关键模块:(1) 频率-空间全向移位(FSO-Shift),对低频分支进行离散小波分解后实施全向空间移位,强化全局上下文表征同时保留高频解剖细节;(2) 结构保真增强块(SFEB),通过频域感知特征融合自适应强化解剖结构。在UNC和BNU数据集上的综合实验表明,FS-RWKV在T1w和T2w模态下均持续超越现有基于CNN、Transformer、GAN及RWKV的基线方法,显著提升解剖保真度与视觉感知质量。
原文摘要 · Abstract (English)
Ultra-high-field 7T MRI offers enhanced spatial resolution and tissue contrast that enables the detection of subtle pathological changes in neurological disorders. However, the limited availability of 7T scanners restricts widespread clinical adoption due to substantial infrastructure costs and technical demands. Computational approaches for synthesizing 7T-quality images from accessible 3T acquisitions present a viable solution to this accessibility challenge. Existing CNN approaches suffer from limited spatial coverage, while Transformer models demand excessive computational overhead. RWKV architectures offer an efficient alternative for global feature modeling in medical image synthesis, combining linear computational complexity with strong long-range dependency capture. Building on this foundation, we propose Frequency Spatial-RWKV (FS-RWKV), an RWKV-based framework for 3T-to-7T MRI translation. To better address the challenges of anatomical detail preservation and global tissue contrast recovery, FS-RWKV incorporates two key modules: (1) Frequency-Spatial Omnidirectional Shift (FSO-Shift), which performs discrete wavelet decomposition followed by omnidirectional spatial shifting on the low-frequency branch to enhance global contextual representation while preserving high-frequency anatomical details; and (2) Structural Fidelity Enhancement Block (SFEB), a module that adaptively reinforces anatomical structure through frequency-aware feature fusion. Comprehensive experiments on UNC and BNU datasets demonstrate that FS-RWKV consistently outperforms existing CNN-, Transformer-, GAN-, and RWKV-based baselines across both T1w and T2w modalities, achieving superior anatomical fidelity and perceptual quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。