arXiv:2506.22567cs.CVcs.AI2025-06被引 19

用多模型知识蒸馏,打造通用医学视觉语言模型

Unifying Biomedical Vision-Language Expertise: Towards a Generalist Foundation Model via Multi-CLIP Knowledge Distillation

  • 从9个医学CLIP模型蒸馏知识,避免依赖海量原始数据
  • 在58个数据集上表现超越所有教师模型,泛化能力强
  • 适合医学视觉分析、跨模态检索等研究者使用

基于自然图像的百亿级图文对预训练的CLIP模型在零样本分类、跨模态检索和开放问答任务中表现出色。然而,将其成功迁移至生物医学领域受限于大规模生物医学图文语料稀缺、图像模态异质性以及机构间数据标准不统一。为此,我们提出MMKD-CLIP,通过多医学CLIP知识蒸馏构建通用生物医学基础模型。该模型不依赖百亿级原始数据,而是从9个先进领域专用或通用医学CLIP模型中蒸馏知识,每个模型均在数百万医学图文对上预训练。采用两阶段训练:第一阶段在26种图像模态的超过290万医学图文对上进行CLIP式预训练;第二阶段利用1920万特征对进行特征级蒸馏。在涵盖9种图像模态、超1080万医学图像的58个多样化生物医学数据集上评估,覆盖零样本分类、线性探测、跨模态检索、视觉问答、生存预测和癌症诊断六类核心任务。MMKD-CLIP始终优于所有教师模型,并在不同图像域和任务设置下展现出卓越鲁棒性与泛化能力。结果表明,多教师知识蒸馏是应对真实世界数据限制下构建高性能生物医学基础模型的有效可扩展范式。

原文摘要 · Abstract (English)

CLIP models pretrained on natural images with billion-scale image-text pairs have demonstrated impressive capabilities in zero-shot classification, cross-modal retrieval, and open-ended visual answering. However, transferring this success to biomedicine is hindered by the scarcity of large-scale biomedical image-text corpora, the heterogeneity of image modalities, and fragmented data standards across institutions. These limitations hinder the development of a unified and generalizable biomedical foundation model trained from scratch. To overcome this, we introduce MMKD-CLIP, a generalist biomedical foundation model developed via Multiple Medical CLIP Knowledge Distillation. Rather than relying on billion-scale raw data, MMKD-CLIP distills knowledge from nine state-of-the-art domain-specific or generalist biomedical CLIP models, each pretrained on millions of biomedical image-text pairs. Our two-stage training pipeline first performs CLIP-style pretraining on over 2.9 million biomedical image-text pairs from 26 image modalities, followed by feature-level distillation using over 19.2 million feature pairs extracted from teacher models. We evaluate MMKD-CLIP on 58 diverse biomedical datasets, encompassing over 10.8 million biomedical images across nine image modalities. The evaluation spans six core task types: zero-shot classification, linear probing, cross-modal retrieval, visual question answering, survival prediction, and cancer diagnosis. MMKD-CLIP consistently outperforms all teacher models while demonstrating remarkable robustness and generalization across image domains and task settings. These results underscore that multi-teacher knowledge distillation is a scalable and effective paradigm for building high-performing biomedical foundation models under the practical constraints of real-world data availability.

医学视觉知识蒸馏基础模型跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。