提出可处理任意数量模态的医学图像融合方法,提升临床实用性。
FlexiD-Fuse: Flexible number of inputs multi-modal medical image fusion based on diffusion model
- 基于扩散模型与贝叶斯推断,将融合问题转化为最大似然估计。
- 支持二模态与三模态输入,端到端生成高质量跨模态融合图像。
- 适用于多模态医学图像及红外可见光等多场景融合,适合临床部署。
不同模态的医学影像可提供疾病独特的生理与解剖信息。多模态医学图像融合通过整合互补影像中的有用信息,生成全面反映病灶特征的融合图像,辅助医生临床诊断。然而,现有融合方法仅能处理固定数量的模态输入(如仅支持双模态或三模态),无法直接处理输入数量可变的情况,限制了其在临床中的应用。为此,我们提出 FlexiD-Fuse,一种基于扩散模型的图像融合网络,可灵活处理任意数量的输入模态。该方法将仅支持固定条件输入的扩散融合问题,转化为基于扩散过程和分层贝叶斯建模的最大似然估计问题。通过在扩散采样迭代中引入期望最大化算法,FlexiD-Fuse 能够从源图像中生成包含跨模态信息的高质量融合图像,且不依赖输入图像数量。我们在哈佛数据集上对比最新双模态与三模态融合方法,并使用九个主流指标进行评估,结果表明本方法在可变输入下表现最优。此外,我们还在红外-可见光、多曝光、多焦点等任务中进行了扩展实验,验证了方法在任意数量输入下的有效性与优越性。
原文摘要 · Abstract (English)
Different modalities of medical images provide unique physiological and anatomical information for diseases. Multi-modal medical image fusion integrates useful information from different complementary medical images with different modalities, producing a fused image that comprehensively and objectively reflects lesion characteristics to assist doctors in clinical diagnosis. However, existing fusion methods can only handle a fixed number of modality inputs, such as accepting only two-modal or tri-modal inputs, and cannot directly process varying input quantities, which hinders their application in clinical settings. To tackle this issue, we introduce FlexiD-Fuse, a diffusion-based image fusion network designed to accommodate flexible quantities of input modalities. It can end-to-end process two-modal and tri-modal medical image fusion under the same weight. FlexiD-Fuse transforms the diffusion fusion problem, which supports only fixed-condition inputs, into a maximum likelihood estimation problem based on the diffusion process and hierarchical Bayesian modeling. By incorporating the Expectation-Maximization algorithm into the diffusion sampling iteration process, FlexiD-Fuse can generate high-quality fused images with cross-modal information from source images, independently of the number of input images. We compared the latest two and tri-modal medical image fusion methods, tested them on Harvard datasets, and evaluated them using nine popular metrics. The experimental results show that our method achieves the best performance in medical image fusion with varying inputs. Meanwhile, we conducted extensive extension experiments on infrared-visible, multi-exposure, and multi-focus image fusion tasks with arbitrary numbers, and compared them with the perspective SOTA methods. The results of the extension experiments consistently demonstrate the effectiveness and superiority of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。