构建多时序胸片数据集,提升模型对疾病进展的判断能力
MMXU: A Multi-Modal and Multi-X-ray Understanding Dataset for Disease Progression
- 设计多时序影像对比任务,融合历史病历与当前影像
- 引入历史记录增强后,诊断准确率提升至少20%
- 适合研究医学图像时序分析与大模型临床应用者
大型视觉语言模型在医疗视觉问答和图像诊断中展现出巨大潜力,但现有数据集和模型常忽视病史整合与疾病进展分析等关键因素。本文提出MMXU(多模态多胸片理解)数据集,专攻患者两次就诊间特定区域变化的识别。不同于以往仅支持单图问答的数据集,MMXU支持多图对比问题,融合当前与历史影像。实验表明,现有LVLM在MMXU测试集上表现有限,即便在传统基准上表现优异亦然。为此,我们提出医患记录增强生成(MAG)方法,结合全局与局部历史信息。实验显示,引入历史记录使诊断准确率提升至少20%,显著缩小模型与人类专家性能差距。在MMXU-Dev上微调模型也取得明显改进。本工作强调历史上下文对医学图像解读的重要性,推动LVLM在临床诊断中的应用。数据集已开源:https://github.com/linjiemu/MMXU。
原文摘要 · Abstract (English)
Large vision-language models (LVLMs) have shown great promise in medical applications, particularly in visual question answering (MedVQA) and diagnosis from medical images. However, existing datasets and models often fail to consider critical aspects of medical diagnostics, such as the integration of historical records and the analysis of disease progression over time. In this paper, we introduce MMXU (Multimodal and MultiX-ray Understanding), a novel dataset for MedVQA that focuses on identifying changes in specific regions between two patient visits. Unlike previous datasets that primarily address single-image questions, MMXU enables multi-image questions, incorporating both current and historical patient data. We demonstrate the limitations of current LVLMs in identifying disease progression on MMXU-\textit{test}, even those that perform well on traditional benchmarks. To address this, we propose a MedRecord-Augmented Generation (MAG) approach, incorporating both global and regional historical records. Our experiments show that integrating historical records significantly enhances diagnostic accuracy by at least 20\%, bridging the gap between current LVLMs and human expert performance. Additionally, we fine-tune models with MAG on MMXU-\textit{dev}, which demonstrates notable improvements. We hope this work could illuminate the avenue of advancing the use of LVLMs in medical diagnostics by emphasizing the importance of historical context in interpreting medical images. Our dataset is released at github: https://github.com/linjiemu/MMXU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。