用大模型引导诊断证据对齐,提升有限配对数据下的医学图文预训练效果。
LLM-Guided Diagnostic Evidence Alignment for Medical Vision-Language Pretraining under Limited Pairing
- 利用大模型提取影像报告中的关键诊断证据,构建共享证据空间。
- 在仅有少量配对数据下,实现比传统方法更优的图文对齐与零样本分类性能。
- 适合医学视觉语言模型在标注数据稀缺场景下的高效预训练。
现有基于CLIP的医学视觉-语言预训练方法依赖大量配对数据进行全局或局部对齐,但全局对齐易受非诊断信息干扰,局部对齐又难以整合关键诊断证据,导致难以学习可靠的诊断表征,限制了其在配对数据有限的医疗场景中的应用。为此,本文提出一种大模型引导的诊断证据对齐方法(LGDEA),将预训练目标转向更符合临床诊断流程的证据级对齐。具体而言,利用大模型从放射科报告中提取关键诊断证据,构建共享诊断证据空间,实现感知证据的跨模态对齐,使LGDEA能有效利用大量未配对的医学图像与报告,显著降低对配对数据的依赖。大量实验表明,该方法在短语定位、图文检索和零样本分类任务上均取得持续且显著的提升,甚至媲美依赖大量配对数据的预训练方法。
原文摘要 · Abstract (English)
Most existing CLIP-style medical vision--language pretraining methods rely on global or local alignment with substantial paired data. However, global alignment is easily dominated by non-diagnostic information, while local alignment fails to integrate key diagnostic evidence. As a result, learning reliable diagnostic representations becomes difficult, which limits their applicability in medical scenarios with limited paired data. To address this issue, we propose an LLM-Guided Diagnostic Evidence Alignment method (LGDEA), which shifts the pretraining objective toward evidence-level alignment that is more consistent with the medical diagnostic process. Specifically, we leverage LLMs to extract key diagnostic evidence from radiology reports and construct a shared diagnostic evidence space, enabling evidence-aware cross-modal alignment and allowing LGDEA to effectively exploit abundant unpaired medical images and reports, thereby substantially alleviating the reliance on paired data. Extensive experimental results demonstrate that our method achieves consistent and significant improvements on phrase grounding, image--text retrieval, and zero-shot classification, and even rivals pretraining methods that rely on substantial paired data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。