arXiv:2604.08936cs.CV2026-04

医学大模型通过分解信息,让不同模态图像特征更清晰、多样。

M-IDoL: Information Decomposition for Modality-Specific and Diverse Representation Learning in Medical Foundation Model

论文配图:M-IDoL: Information Decomposition for Modality-Specific and Diverse Representation Learning in Medical Foundation Model
图 1 · 摘自论文原文
  • 用混合专家结构分离多模态特征,提升模态特异性
  • 在单模态内精细区分语义,增强特征多样性
  • 适合需要精准医学图像分析的研究者使用

医学基础模型(MFMs)旨在从多模态医学图像中学习通用表示,以泛化到多样下游临床任务。然而,现有模型普遍存在信息模糊问题,导致多模态表示在单一嵌入空间中混合,削弱了模态特异性和多样性。本文提出M-IDoL,一种自监督医学基础模型,通过信息分解实现多模态表征学习:一是最大化模态间熵,将多模态表示分散至可分的混合专家(MoE)子空间,实现模态间的表征特异性;二是最小化模态内不确定性,通过在每个MoE子空间内进行细粒度语义区分,丰富单模态表征多样性。在115万张医学图像上预训练后,M-IDoL在21个下游临床任务中表现卓越,优于20个基础模型,在五种成像模态(如X光、眼底、OCT、皮肤镜和病理)上均取得领先,并展现出更清晰的模态间特征聚类分离与模态内更细粒度的特征区分能力。

原文摘要 · Abstract (English)

Medical foundation models (MFMs) aim to learn universal representations from multimodal medical images that can generalize effectively to diverse downstream clinical tasks. However, most existing MFMs suffer from information ambiguity that blends multimodal representations in a single embedding space, leading to the degradation of modality specificity and diversity. In this paper, we propose M-IDoL, a self-supervised MFM that introduces Information Decomposition for multimodal representation Learning via two objectives: i) maximizing inter-modality entropy by dispersing multimodal representations into separable Mixture-of-Experts (MoE) subspaces to achieve representation specificity across modalities; and ii) minimizing intra-modality uncertainty by performing fine-grained semantic discrimination within each MoE subspace to enrich representation diversity per modality. By pre-training on 1.15 million medical images, M-IDoL i) delivers superior generalization across 21 downstream clinical tasks, outperforming 20 foundation models on five imaging modalities (e.g., X-ray, fundus, OCT, dermoscopy and pathology), and ii) learns modality-specific and diverse representations, showing clearer separation of feature clusters across modalities and finer-grained feature discrimination within each modality.

医学图像多模态表征学习MoE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。