arXiv:2505.12251cs.CV2025-05被引 6

用医学先验知识融合多模态影像,提升诊断信息保真度。

SMFusion: Semantic-Preserving Fusion of Multimodal Medical Images for Enhanced Clinical Diagnosis

  • 引入医学文本先验,通过语义对齐模块融合图文特征。
  • 在公开数据集上实现更优的图像质量与诊断报告生成效果。
  • 适合临床辅助诊断系统开发人员参考使用。

多模态医学图像融合通过整合不同模态的互补信息,提升图像可读性与临床应用价值。然而,现有方法主要遵循计算机视觉范式进行特征提取与融合策略设计,忽视了医学图像中蕴含的丰富语义信息。为此,本文提出一种新颖的语义引导医学图像融合方法,首次将医学先验知识融入融合过程。具体地,构建了一个公开的多模态医学图像-文本数据集,利用BiomedGPT生成文本描述,并通过语义交互对齐模块在高维空间中将文本特征与图像特征进行语义对齐。该过程采用基于交叉注意力的线性变换,自动映射文本与视觉特征间的关系,促进全面学习。对齐后的特征被输入文本注入模块进行特征级融合。不同于传统方法,本文进一步从融合图像生成诊断报告,以评估医学信息保留程度。同时,设计医疗语义损失函数以增强源图像中文本线索的保留。在测试数据集上的实验结果表明,所提方法在定性和定量评估中均表现更优,且保留了更多关键医学信息。

原文摘要 · Abstract (English)

Multimodal medical image fusion plays a crucial role in medical diagnosis by integrating complementary information from different modalities to enhance image readability and clinical applicability. However, existing methods mainly follow computer vision standards for feature extraction and fusion strategy formulation, overlooking the rich semantic information inherent in medical images. To address this limitation, we propose a novel semantic-guided medical image fusion approach that, for the first time, incorporates medical prior knowledge into the fusion process. Specifically, we construct a publicly available multimodal medical image-text dataset, upon which text descriptions generated by BiomedGPT are encoded and semantically aligned with image features in a high-dimensional space via a semantic interaction alignment module. During this process, a cross attention based linear transformation automatically maps the relationship between textual and visual features to facilitate comprehensive learning. The aligned features are then embedded into a text-injection module for further feature-level fusion. Unlike traditional methods, we further generate diagnostic reports from the fused images to assess the preservation of medical information. Additionally, we design a medical semantic loss function to enhance the retention of textual cues from the source images. Experimental results on test datasets demonstrate that the proposed method achieves superior performance in both qualitative and quantitative evaluations while preserving more critical medical information.

医学影像多模态融合语义对齐诊断报告

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。