arXiv:2609.04781cs.CV2026-09

用协作门控MLP实现医学图像分割中的细粒度跨模态融合。

CoMLP: Cooperatively-Gated MLPs for Fine-Grained Cross-Modal Information Fusion in Medical Image Segmentation

论文配图:CoMLP: Cooperatively-Gated MLPs for Fine-Grained Cross-Modal Information Fusion in Medical Image Segmentation
图 1 · 摘自论文原文
  • 通过互补的局部与全局MLP交互建模跨模态依赖关系。
  • 在5个基准上优于当前最佳方法,尤其在高分辨率特征图上表现优异。
  • 适合需要融合多模态医学影像与临床报告的场景。

多模态医学影像与临床报告提供了互补的解剖、功能和语义信息,用于医学图像分割。有效利用这些异构源需要精细的跨模态信息融合,既要保留细微的空间细节,又要捕捉模态间的语义依赖。现有融合方法常依赖交叉注意力,其计算开销随空间分辨率快速增加,使得在高分辨率特征图上进行密集跨模态交互变得困难,尤其对于体积医学图像。本文提出CoMLP,一种用于医学图像分割中细粒度跨模态信息融合的协作门控MLP模块。CoMLP通过协作式交叉门控建模跨模态依赖,基于互补的区域和扩张MLP交互,以捕捉局部与全局跨模态依赖。我们进一步设计了多源融合架构,其中CoMLP同时执行跨成像模态的图像间融合及视觉特征与文本报告之间的视觉-语言融合,无需依赖密集交叉注意力即可整合异构信息。在涵盖2D/3D图像、临床报告、多种成像模态和不同解剖区域的五个医学分割基准上的大量实验表明,该方法在性能上持续优于当前最先进的多模态与语言引导分割方法。消融实验进一步表明,高空间分辨率下的细粒度交互以及互补的局部-全局融合是性能提升的关键。这些结果展示了基于MLP的交互作为医学图像分割中细粒度跨模态信息融合的有效替代方案的潜力。

原文摘要 · Abstract (English)

Multi-modal medical images and clinical reports provide complementary anatomical, functional, and semantic information for medical image segmentation. Effectively exploiting these heterogeneous sources requires fine-grained cross-modal information fusion that preserves subtle spatial details while capturing semantic dependencies across modalities. Existing fusion approaches frequently rely on cross-attention, whose computational burden increases rapidly with spatial resolution, making dense cross-modal interaction difficult on high-resolution feature maps, particularly for volumetric medical images. In this work, we propose CoMLP, a cooperatively-gated MLP module for fine-grained cross-modal information fusion in medical image segmentation. CoMLP models cross-modal dependencies through cooperative cross-gating, built upon complementary regional and dilated MLP interactions, to capture local and global cross-modal dependencies. We further develop a multi-source fusion architecture in which CoMLP performs both inter-image fusion across imaging modalities and vision-language fusion between visual features and textual reports, enabling heterogeneous information to be integrated without relying on dense cross-attention. Extensive experiments on five medical segmentation benchmarks, covering 2D/3D images, clinical reports, multiple imaging modalities, and diverse anatomical regions, demonstrate consistent improvements over state-of-the-art multi-modal and language-guided segmentation methods. Ablation studies further show that fine-grained interaction at high spatial resolutions and complementary local-global fusion are critical to the performance gains. These results demonstrate the potential of MLP-based interaction as an effective alternative for fine-grained cross-modal information fusion in medical image segmentation.

医学图像跨模态融合MLP分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。