联合处理多模态医学图像融合与超分辨率,提升诊断清晰度。
TriFusion-SR: Joint Tri-Modal Medical Image Fusion and SR
- 用小波变换分解多模态特征,实现频域感知的跨模态交互。
- 在多个上采样尺度下,PSNR提升4.8%-12.4%,RMSE和LPIPS显著降低。
- 适合医学影像分析、精准诊断领域研究人员使用。
多模态医学图像融合通过整合结构与功能信息促进全面诊断,但受限于分辨率下降和模态差异。现有方法通常分步进行图像融合与超分辨率(SR),导致伪影和感知质量下降。这一问题在结合解剖模态(如MRI、CT)与功能扫描(如PET、SPECT)的三模态设置中尤为突出,因频域不平衡更严重。我们提出TriFusionSR,一种基于小波引导的条件扩散框架,实现三模态融合与超分辨率的联合建模。该框架利用二维离散小波变换(2D DWT)显式分解多模态特征为频率带,支持频域感知的跨模态交互。进一步引入修正小波特征(RWF)策略校准潜在系数,并设计带有门控通道-空间注意力的自适应时空融合(ASFF)模块,实现结构驱动的多模态优化。大量实验表明,该方法达到业界领先性能,在多个上采样尺度下实现4.8%-12.4%的PSNR提升,同时大幅降低RMSE与LPIPS。
原文摘要 · Abstract (English)
Multimodal medical image fusion facilitates comprehensive diagnosis by aggregating complementary structural and functional information, but its effectiveness is limited by resolution degradation and modality discrepancies. Existing approaches typically perform image fusion and super-resolution (SR) in separate stages, leading to artifacts and degraded perceptual quality. These limitations are further amplified in tri-modal settings that combine anatomical modalities (e.g., MRI, CT) with functional scans (e.g., PET, SPECT) due to pronounced frequency domain imbalances. We propose TriFusionSR, a wavelet-guided conditional diffusion framework for joint tri-modal fusion and SR. The framework explicitly decomposes multimodal features into frequency bands using the 2D Discrete Wavelet Transform, enabling frequency-aware crossmodal interaction. We further introduce a Rectified Wavelet Features (RWF) strategy for latent coefficient calibration, followed by an Adaptive Spatial-Frequency Fusion (ASFF) module with gated channel-spatial attention to enable structure-driven multimodal refinement. Extensive experiments demonstrate state-of-the-art performance, achieving 4.8-12.4% PSNR improvement and substantial reductions in RMSE and LPIPS across multiple upsampling scales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。