用医学知识增强提示,让影像诊断模型更抗提示变化。
KEPIL: Knowledge-Enhanced Prompt-Image Learning for Prompt-Robust Disease Detection

- 用医学本体和大模型动态优化提示词
- 在七项测试中零样本性能达顶尖水平
- 适合需要稳定诊断的临床部署场景
视觉-语言模型(VLM)在放射科临床决策支持中具有潜力,因其可联合推理医学图像与文本信息,从而利用互补的临床数据。然而实际中放射学发现呈长尾分布,部分病种样本不足,零样本推理至关重要。现有基于CLIP的医学VLM对提示词变化敏感,且推理时缺乏可信外部知识,限制了其临床可靠性。本文提出KEPIL框架,通过整合结构化医学知识提升零样本泛化稳定性。该框架包含:(i) 基于本体与大模型辅助的动态提示增强;(ii) 通过双嵌入目标对齐等效提示变体的语义感知对比损失;(iii) 基于实体的报告标准化以生成本体对齐表示。在七个基准上,KEPIL实现顶尖零样本性能;在提示变异测试中,于CheXpert数据集提升AUC 6.37%,平均提升4.11%。结果表明,结构化知识与鲁棒提示设计是构建可靠放射科VLM的关键。代码将发布于https://github.com/Roypic/KEPIL。
原文摘要 · Abstract (English)
Vision--language models (VLMs) show promise for clinical decision support in radiology because they enable joint reasoning over radiological images and clinical text, thereby leveraging complementary clinical information. However, radiological findings are long-tailed in practice, leaving some conditions underrepresented and making zero-shot inference essential. Yet current CLIP-style medical VLMs are sensitive to prompt variations and often lack trustworthy external knowledge at inference time, which hinders reliable clinical deployment. We present \textit{KEPIL}, a prompt-robust framework that integrates curated medical knowledge to stabilize zero-shot generalization. KEPIL comprises: (i) \emph{dynamic prompt enrichment} using ontologies with LLM assistance, (ii) a \emph{semantic-aware contrastive loss} aligning embeddings of equivalent prompt variants via a dual-embedding objective, and (iii) \emph{entity-centric report standardization} to yield ontology-aligned representations. Across seven benchmarks, KEPIL achieves state-of-the-art zero-shot inference performance; under prompt-variation tests, it improves AUC by \(6.37\%\) on \textit{CheXpert} and by \(4.11\%\) on average. These results suggest that structured knowledge and robust prompt design are key to clinically reliable radiology-facing VLMs. Code will be released at https://github.com/Roypic/KEPIL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。