用频域融合提升医学影像分割准确率
Frequency-domain Multi-modal Fusion for Language-guided Medical Image Segmentation
- 在解码器中通过频域特征双向交互增强视觉表示
- 引入语言引导的频域特征交互模块,抑制无关视觉信息
- 在两个数据集上均优于现有方法,尤其对复杂病灶有效
自动分割放射影像中的感染区域对诊断肺部感染性疾病至关重要。近期研究证明,结合临床文本报告作为语义引导可提升医学图像分割精度。然而,病灶复杂的形态变化以及视觉与语言模态间的固有语义鸿沟,导致现有方法难以有效增强视觉特征表示并消除语义无关信息,最终影响分割性能。为此,我们提出频率域多模态交互模型(FMISeg),用于语言引导的医学图像分割。该模型为后融合架构,在解码器中建立语言特征与频域视觉特征之间的交互。具体地,为增强视觉表示,提出频域特征双向交互(FFBI)模块以高效融合频域特征;同时,在解码器中引入语言引导的频域特征交互(LFFI)模块,借助语言信息抑制语义无关的视觉特征。在QaTa-COV19和MosMedData+数据集上的实验表明,所提方法在定性和定量上均优于当前最优方法。
原文摘要 · Abstract (English)
Automatically segmenting infected areas in radiological images is essential for diagnosing pulmonary infectious diseases. Recent studies have demonstrated that the accuracy of the medical image segmentation can be improved by incorporating clinical text reports as semantic guidance. However, the complex morphological changes of lesions and the inherent semantic gap between vision-language modalities prevent existing methods from effectively enhancing the representation of visual features and eliminating semantically irrelevant information, ultimately resulting in suboptimal segmentation performance. To address these problems, we propose a Frequency-domain Multi-modal Interaction model (FMISeg) for language-guided medical image segmentation. FMISeg is a late fusion model that establishes interaction between linguistic features and frequency-domain visual features in the decoder. Specifically, to enhance the visual representation, our method introduces a Frequency-domain Feature Bidirectional Interaction (FFBI) module to effectively fuse frequency-domain features. Furthermore, a Language-guided Frequency-domain Feature Interaction (LFFI) module is incorporated within the decoder to suppress semantically irrelevant visual features under the guidance of linguistic information. Experiments on QaTa-COV19 and MosMedData+ demonstrated that our method outperforms the state-of-the-art methods qualitatively and quantitatively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。