arXiv:2409.00341cs.CV2024-09被引 18

用大模型知识指导医学影像分析,提升诊断准确率。

Aligning Medical Images with General Knowledge from Large Language Models

  • 通过视觉症状生成器提取可解释的医学特征
  • 双提示网络在两个数据集上超越现有方法
  • 适合医疗影像、多模态学习研究者参考

预训练的大规模视觉-语言模型(如CLIP)利用自然语言作为监督信号,革新了视觉表征学习,并展现出良好的泛化能力。本文提出ViP,一种新的视觉症状引导提示学习框架,用于医学图像分析,促进从CLIP中迁移通用知识。ViP包含两个关键组件:视觉症状生成器(VSG)和双提示网络。其中,VSG旨在从预训练的大语言模型中提取可解释的视觉症状;双提示网络则利用这些视觉症状,指导两个可学习提示模块(上下文提示与融合提示)的训练,从而有效适应基于大型视觉-语言模型的医学图像分析。大量实验证明,ViP在两个具有挑战性的数据集上均优于当前最优方法。

原文摘要 · Abstract (English)

Pre-trained large vision-language models (VLMs) like CLIP have revolutionized visual representation learning using natural language as supervisions, and demonstrated promising generalization ability. In this work, we propose ViP, a novel visual symptom-guided prompt learning framework for medical image analysis, which facilitates general knowledge transfer from CLIP. ViP consists of two key components: a visual symptom generator (VSG) and a dual-prompt network. Specifically, VSG aims to extract explicable visual symptoms from pre-trained large language models, while the dual-prompt network utilizes these visual symptoms to guide the training on two learnable prompt modules, i.e., context prompt and merge prompt, which effectively adapts our framework to medical image analysis via large VLMs. Extensive experimental results demonstrate that ViP can outperform state-of-the-art methods on two challenging datasets.

医学影像提示学习大模型跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。