arXiv:2602.06184cs.CVcs.CL2026-02

将医学表型本体知识融入视觉语言模型,提升医疗图像理解的准确性与可解释性。

PhenoLIP: Integrating Phenotype Ontology Knowledge into Medical Vision-Language Pretraining

  • 构建表型中心的多模态知识图谱PhenoKG,包含52万+图文对和3000+表型
  • 通过两阶段知识蒸馏,使模型在表型分类上比BiomedCLIP高8.85%准确率
  • 适合医疗AI研究者和需要可解释性诊断模型的临床应用

近年来,基于CLIP的大型视觉语言模型(VLM)在医疗图像分析中取得显著进展。然而,现有医疗VLM大多依赖粗粒度的图像-文本对比目标,难以捕捉医学表型本体中系统化的视觉知识。为此,我们构建了首个大规模、以表型为中心的多模态知识图谱PhenoKG,包含超过52万条高质量图像-文本对,关联3000多个表型。在此基础上,提出PhenoLIP预训练框架,通过两阶段过程将结构化表型知识融入医疗VLM:首先从文本本体数据中学习增强的表型嵌入空间,再通过教师引导的知识蒸馏目标,将该知识注入多模态预训练。为支持评估,进一步引入PhenoBench——一个专家验证的基准,涵盖7800余张图像-标题对,覆盖1000多个表型。大量实验表明,PhenoLIP优于现有最优基线,在表型分类上比BiomedCLIP提升8.85%,在跨模态检索上比BIOMEDICA提升15.03%,验证了将表型中心先验知识融入医疗VLM对于实现结构化、可解释的医学图像理解的价值。

原文摘要 · Abstract (English)

Recent progress in large-scale CLIP-like vision-language models(VLMs) has greatly advanced medical image analysis. However, most existing medical VLMs still rely on coarse image-text contrastive objectives and fail to capture the systematic visual knowledge encoded in well-defined medical phenotype ontologies. To address this gap, we construct PhenoKG, the first large-scale, phenotype-centric multimodal knowledge graph that encompasses over 520K high-quality image-text pairs linked to more than 3,000 phenotypes. Building upon PhenoKG, we propose PhenoLIP, a novel pretraining framework that explicitly incorporates structured phenotype knowledge into medical VLMs through a two-stage process. We first learn a knowledge-enhanced phenotype embedding space from textual ontology data and then distill this structured knowledge into multimodal pretraining via a teacher-guided knowledge distillation objective. To support evaluation, we further introduce PhenoBench, an expert-verified benchmark designed for phenotype recognition, comprising over 7,800 image--caption pairs covering more than 1,000 phenotypes. Extensive experiments demonstrate that PhenoLIP outperforms previous state-of-the-art baselines, improving upon BiomedCLIP in phenotype classification accuracy by 8.85\% and BIOMEDICA in cross-modal retrieval by 15.03%, underscoring the value of integrating phenotype-centric priors into medical VLMs for structured and interpretable medical image understanding.

视觉语言模型医疗AI知识图谱表型识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。