arXiv:2506.18072eess.IVcs.AI2025-06被引 6

无需配对数据,用共享文本空间实现多模态医学影像对齐

Multimodal Medical Image Binding via Shared Text Embeddings

  • 通过共享文本嵌入空间,将不同医学影像模态自然对齐
  • 在零样本和少样本分类任务上优于现有CLIP类模型
  • 适用于缺乏成对数据的医疗影像分析场景

医学图像分析日益依赖多种成像模态融合以获取互补的解剖与功能信息,从而提升诊断与治疗规划精度。实现跨模态特征对齐是有效多模态分析的关键。尽管对比语言-图像预训练(CLIP)及其变体实现了图像-文本对齐,但它们需要任意两模态间的显式配对数据,在医疗场景中难以获得。为此,我们提出多模态医学图像绑定文本(M³Bind)框架,通过共享文本表示空间,无需任何两模态间的显式配对数据,即可实现多模态医学影像的无缝对齐。具体而言,基于不同图像可自然关联文本的洞察,M³Bind首先微调预训练的CLIP类图像-文本模型,使其模态特异性文本嵌入空间对齐,同时保留原始图像-文本对齐;随后,将各模态特异性文本编码器蒸馏为统一模型,构建共享文本嵌入空间。在X-ray、CT、眼底、心电图及病理图像上的多个下游任务实验表明,相比其CLIP类基线,M³Bind在零样本、少样本分类与跨模态检索任务中均达到最先进性能。结果验证了M³Bind在医疗分析中实现跨图像模态对齐的有效性。

原文摘要 · Abstract (English)

Medical image analysis increasingly relies on the integration of multiple imaging modalities to capture complementary anatomical and functional information, enabling more accurate diagnosis and treatment planning. Achieving aligned feature representations across these diverse modalities is therefore important for effective multimodal analysis. While contrastive language-image pre-training (CLIP) and its variant have enabled image-text alignments, they require explicitly paired data between arbitrary two modalities, which is difficult to acquire in medical contexts. To address the gap, we present Multimodal Medical Image Binding with Text (M\textsuperscript{3}Bind), a novel pre-training framework that enables seamless alignment of multiple medical imaging modalities through a shared text representation space without requiring explicit paired data between any two medical image modalities. Specifically, based on the insight that different images can naturally bind with text, M\textsuperscript{3}Bind first fine-tunes pre-trained CLIP-like image-text models to align their modality-specific text embedding space while preserving their original image-text alignments. Subsequently, we distill these modality-specific text encoders into a unified model, creating a shared text embedding space. Experiments on X-ray, CT, retina, ECG, and pathological images on multiple downstream tasks demonstrate that M\textsuperscript{3}Bind achieves state-of-the-art performance in zero-shot, few-shot classification and cross-modal retrieval tasks compared to its CLIP-like counterparts. These results validate M\textsuperscript{3}Bind's effectiveness in achieving cross-image-modal alignment for medical analysis.

多模态医学影像图像对齐文本嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。