LUMEN通过多模态指令微调,提升纵向胸片的诊断与预后分析能力。
LUMEN: Longitudinal Multi-Modal Radiology Model for Prognosis and Diagnosis
- 基于多图像多任务指令微调,专为纵向胸片设计
- 在MIMIC-CXR数据集上诊断性能显著优于基线模型
- 构建新指令数据集,支持临床预后预测任务
大型视觉语言模型(VLM)已从通用应用转向医疗领域,展现出辅助放射科医生决策的潜力。通过视觉与自然语言问答接口分析胸片(CXR)数据,可辅助诊断。当有纵向影像时,放射科医生需分析时间变化以准确诊断和预后。手动分析耗时,促使开发具备预后能力的训练框架。本文提出LUMEN框架,专为纵向胸片解读优化,采用多图像多任务指令微调,提升诊断与预后性能。在公开数据集MIMIC-CXR及其关联的Medical-Diff-VQA上进行实验,并构建包含纵向研究的新指令跟随数据集,支持预后问答任务。结果表明,该方法在诊断问答任务中显著优于基线模型,且展现出良好的预后预测潜力。这些成果凸显了精心设计的指令微调VLM在提升纵向影像临床解读准确性方面的价值。
原文摘要 · Abstract (English)
Large vision-language models (VLMs) have evolved from general-purpose applications to specialized use cases such as in the clinical domain, demonstrating potential for decision support in radiology. One promising application is assisting radiologists in decision-making by the analysis of radiology imaging data such as chest X-rays (CXR) via a visual and natural language question-answering (VQA) interface. When longitudinal imaging is available, radiologists analyze temporal changes, which are essential for accurate diagnosis and prognosis. The manual longitudinal analysis is a time-consuming process, motivating the development of a training framework that can provide prognostic capabilities. We introduce a novel training framework LUMEN, that is optimized for longitudinal CXR interpretation, leveraging multi-image and multi-task instruction fine-tuning to enhance prognostic and diagnostic performance. We conduct experiments on the publicly available MIMIC-CXR and its associated Medical-Diff-VQA datasets. We further formulate and construct a novel instruction-following dataset incorporating longitudinal studies, enabling the development of a prognostic VQA task. Our method demonstrates significant improvements over baseline models in diagnostic VQA tasks, and more importantly, shows promising potential for prognostic capabilities. These results underscore the value of well-designed, instruction-tuned VLMs in enabling more accurate and clinically meaningful radiological interpretation of longitudinal radiological imaging data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。