用健康参考图提升医学影像诊断准确率
See-in-Pairs: Reference Image-Guided Comparative Vision-Language Models for Medical Diagnosis
- 引入查询图与健康对照图对比,通过提示词实现比较诊断
- 在多个任务上诊断准确率显著提升,小样本微调可进一步优化
- 适合临床辅助诊断系统开发,尤其关注细微异常识别
医学影像诊断困难,因疾病常与正常解剖结构相似且患者间差异大。临床医生常通过跨患者对比健康对照图像来发现细微但关键的异常。尽管健康参考图在实践中丰富,现有医学视觉语言模型(VLM)多基于单图或单序列,缺乏显式比较机制。本文研究将临床启发的对比引入VLM是否能提升性能。结果表明,同时提供查询图像和匹配的健康参考图像,并辅以跨患者对比提示,可显著提升诊断效果。通过少量数据的轻量级监督微调(SFT),性能进一步增强。我们评估了多种参考图选择策略:随机采样、人口统计特征匹配、基于嵌入的检索和跨中心选择,均表现良好。此外,理论分析显示,对比诊断提升了样本效率并增强了视觉-文本表示对齐。研究证实了基于对比的诊断具有临床意义,提供了实用的参考图集成方法,并在多种医学影像任务中取得改进。
原文摘要 · Abstract (English)
Medical image diagnosis is challenging because many diseases resemble normal anatomy and exhibit substantial interpatient variability. Clinicians routinely rely on comparative diagnosis, such as referencing cross-patient healthy control images to identify subtle but clinically meaningful abnormalities. Although healthy reference images are abundant in practice, existing medical vision-language models (VLMs) primarily operate in a single-image or single-series setting and lack explicit mechanisms for comparative diagnosis. This work investigates whether incorporating clinically motivated comparison can enhance VLM performance. We show that providing VLMs with both a query image and a matched healthy reference image, accompanied by cross-patient comparative prompts, significantly improves diagnostic performance. This performance can be further augmented by lightweight supervised fine-tuning (SFT) on a small amount of data. At the same time, we evaluate multiple strategies for selecting reference images, including random sampling, demographic attribute matching, embedding-based retrieval, and cross-center selection, and find consistently strong performance across all settings. Finally, we investigate why comparative diagnosis is effective theoretically, and observe improved sample efficiency and tighter alignment between visual and textual representations. Our findings highlight the clinical relevance of comparison-based diagnosis, provide practical strategies for incorporating reference images into VLMs, and demonstrate improved performance across diverse medical imaging tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。