arXiv:2503.02917eess.IVcs.AI2025-03中稿 · Information Proces…被引 7

用医学概念引导视觉语言模型,让眼科疾病诊断更准确且可解释。

Interpretable Few-Shot Retinal Disease Diagnosis with Concept-Guided Prompting of Vision-Language Models

  • 用GPT提取眼底病可解释概念,作为提示词训练多模态模型。
  • 16次少样本学习平均精度提升5.8%,零样本检测提升2.7%。
  • 适合需要可解释性诊断的临床场景与医疗AI研究者。

深度学习在利用彩色眼底图像分类视网膜疾病方面展现出巨大潜力。然而,现有方法主要依赖图像数据,诊断过程缺乏可解释性,且将医务人员仅视为标注者。为此,本文提出两种关键策略:利用GPT知识库提取视网膜疾病的可解释概念,并将其作为语言成分融入提示学习中,训练结合眼底图像与对应概念的视觉-语言(VL)模型。该方法不仅提升了视网膜疾病分类性能,还增强了少样本和零样本检测(新疾病识别)能力,同时提供基于概念的模型可解释性。在两个多样化的眼底图像数据集上的广泛评估表明,通过引入概念,基于VL模型的少样本方法获得显著性能提升:16次学习平均精度提升约5.8%,零样本(新类别)检测提升2.7%。该方法为真实临床应用中的可解释、高效视网膜疾病识别迈出了关键一步。

原文摘要 · Abstract (English)

Recent advancements in deep learning have shown significant potential for classifying retinal diseases using color fundus images. However, existing works predominantly rely exclusively on image data, lack interpretability in their diagnostic decisions, and treat medical professionals primarily as annotators for ground truth labeling. To fill this gap, we implement two key strategies: extracting interpretable concepts of retinal diseases using the knowledge base of GPT models and incorporating these concepts as a language component in prompt-learning to train vision-language (VL) models with both fundus images and their associated concepts. Our method not only improves retinal disease classification but also enriches few-shot and zero-shot detection (novel disease detection), while offering the added benefit of concept-based model interpretability. Our extensive evaluation across two diverse retinal fundus image datasets illustrates substantial performance gains in VL-model based few-shot methodologies through our concept integration approach, demonstrating an average improvement of approximately 5.8\% and 2.7\% mean average precision for 16-shot learning and zero-shot (novel class) detection respectively. Our method marks a pivotal step towards interpretable and efficient retinal disease recognition for real-world clinical applications.

医学影像少样本学习可解释AI视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。