arXiv:2607.27806cs.CV2026-07

构建首个纵向医学视觉问答基准,评估模型对疾病进展的时序理解能力。

LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA

论文配图:LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA
图 1 · 摘自论文原文
  • 基于医疗知识图谱与时间演化建模,自动生成20.6万组纵向医学图像问答对。
  • 通用与医学大模型在该任务上表现不佳,揭示其时序推理能力严重不足。
  • 提出MedLong-8B模型并分析失败模式,助力提升医学影像动态分析能力。

临床实践中,患者常经历多次随访影像检查,形成纵向数据。建模此类时间信息对于可靠评估疾病进展和治疗反应至关重要。然而,尽管多模态大语言模型(MLLMs)快速发展,纵向医学视觉推理仍鲜受关注。为此,我们提出LoMeVQA,一个包含20.6万组纵向视觉问答对的综合性基准,涵盖五项任务:进展分类、进展描述、进展报告生成、差异区域定位与描述。通过自动化流程,按时间顺序组织病历,利用医疗知识图谱提取临床实体,并建模其时序演变,以指导大模型生成高质量纵向VQA对。大量评估表明,通用及医学领域MLLM在该基准上表现欠佳,暴露出显著的时序推理缺陷。为应对这一问题,我们提出MedLong-8B,在所有任务中达到当前最优性能。此外,我们进行深入分析,揭示关键失败模式,为改进纵向医学视觉推理提供洞见。数据集已开源:https://github.com/pepperbubble/LoMeVQA。

原文摘要 · Abstract (English)

In clinical practice, patients often undergo multiple imaging examinations over successive visits, yielding longitudinal data. Modeling such temporal information is crucial for reliable assessment of disease progression and treatment response. However, despite the rapid advancement of multimodal large language models (MLLMs), longitudinal medical visual reasoning remains largely underexplored. To fill this gap, we propose LoMeVQA, a comprehensive benchmark consisting of 206K longitudinal visual question answering (VQA) pairs for temporal medical image analysis. LoMeVQA covers five tasks: progress classification, progress description, progress report generation, differential region grounding, and differential region description. To construct the dataset, we develop an automated pipeline that (1) organizes patient records chronologically, (2) extracts clinically meaningful entities via a medical knowledge graph, and (3) models their temporal evolution to guide large language models in generating high-quality longitudinal VQA pairs. Extensive evaluations demonstrate that both general-purpose and medical-domain MLLMs perform poorly on LoMeVQA, revealing substantial limitations in temporal reasoning. To address these limitations, we introduce MedLong-8B, which achieves state-of-the-art performance across all tasks. Beyond benchmarking, we conduct detailed analyses that uncover key failure modes and shed light on how to improve longitudinal medical visual reasoning. Our data is available at: https://github.com/pepperbubble/LoMeVQA

医学视觉问答纵向分析时序推理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。