arXiv:2501.11231cs.CV2025-01AAAI被引 12

无需训练,从CLIP中挖掘医学知识提升图像诊断准确率

KPL: Training-Free Medical Knowledge Mining of Vision-Language Models

  • 用文本代理优化和多模态代理学习增强医学描述
  • 在医疗与自然图像数据集上均超越现有基线方法
  • 适合零样本医学图像分类研究者使用

视觉语言模型如CLIP因大规模图文预训练在图像识别中表现优异。但在零样本医学图像诊断应用中,仅用单一类别名称表示图像类别的局限性及视觉与文本空间间的模态差距导致性能受限。尽管已有研究尝试通过大语言模型丰富疾病描述,但缺乏类别特异性知识仍使效果不佳。此外,现有代理学习方法在自然图像数据集上表现稳定,但在医疗数据集上却存在不稳定性。为此,本文提出知识代理学习(KPL),通过文本代理优化与多模态代理学习,从CLIP中挖掘医学知识。KPL从构建的知识增强基础中检索与图像相关的知识描述,以丰富语义文本代理,并利用输入图像与这些描述经CLIP编码后生成稳定的多模态代理,显著提升零样本分类性能。在医疗与自然图像数据集上的大量实验表明,KPL实现了有效的零样本图像分类,优于所有基线方法。结果凸显了从CLIP中挖掘知识用于医学图像分类的巨大潜力。

原文摘要 · Abstract (English)

Visual Language Models such as CLIP excel in image recognition due to extensive image-text pre-training. However, applying the CLIP inference in zero-shot classification, particularly for medical image diagnosis, faces challenges due to: 1) the inadequacy of representing image classes solely with single category names; 2) the modal gap between the visual and text spaces generated by CLIP encoders. Despite attempts to enrich disease descriptions with large language models, the lack of class-specific knowledge often leads to poor performance. In addition, empirical evidence suggests that existing proxy learning methods for zero-shot image classification on natural image datasets exhibit instability when applied to medical datasets. To tackle these challenges, we introduce the Knowledge Proxy Learning (KPL) to mine knowledge from CLIP. KPL is designed to leverage CLIP's multimodal understandings for medical image classification through Text Proxy Optimization and Multimodal Proxy Learning. Specifically, KPL retrieves image-relevant knowledge descriptions from the constructed knowledge-enhanced base to enrich semantic text proxies. It then harnesses input images and these descriptions, encoded via CLIP, to stably generate multimodal proxies that boost the zero-shot classification performance. Extensive experiments conducted on both medical and natural image datasets demonstrate that KPL enables effective zero-shot image classification, outperforming all baselines. These findings highlight the great potential in this paradigm of mining knowledge from CLIP for medical image classification and broader areas.

医学图像零样本学习CLIP知识挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。