arXiv:2505.19031cs.CVcs.AI2025-05被引 15

构建医疗多图理解数据集,提升大模型对多张医学图像的分析能力。

Medical Large Vision Language Models with Multi-Image Visual Ability

  • 构建83.2万条多图医学问答数据,涵盖时序、推理等四种视觉能力
  • 新模型在多图医学问答上显著超越现有模型,尤其在跨图像推理任务中
  • 适合研究医疗视觉语言模型、多图医学影像分析的学者使用

医学大型视觉语言模型(LVLMs)在单图问答任务中表现良好,但在多图临床场景下的能力仍待探索。与单图任务不同,多图医学任务常需复杂视觉理解,如时序推理和跨模态分析,而现有模型支持不足。为此,我们提出Med-MIM指令数据集,包含83.2万条涵盖时序理解、推理、比较、共指四类多图视觉能力的医学多图问答对。基于该数据集,我们微调Mantis和LLaVA-Med,得到两个专用于多图分析的医学视觉语言模型:MIM-LLaVA-Med和Med-Mantis。同时,我们构建了Med-MIM基准测试,全面评估八种主流LVLM在多图医学理解上的表现。实验表明,两新模型在内部和外部测试集上均取得领先,证明该数据集有效提升了医疗领域模型的多图理解能力。

原文摘要 · Abstract (English)

Medical large vision-language models (LVLMs) have demonstrated promising performance across various single-image question answering (QA) benchmarks, yet their capability in processing multi-image clinical scenarios remains underexplored. Unlike single image based tasks, medical tasks involving multiple images often demand sophisticated visual understanding capabilities, such as temporal reasoning and cross-modal analysis, which are poorly supported by current medical LVLMs. To bridge this critical gap, we present the Med-MIM instruction dataset, comprising 83.2K medical multi-image QA pairs that span four types of multi-image visual abilities (temporal understanding, reasoning, comparison, co-reference). Using this dataset, we fine-tune Mantis and LLaVA-Med, resulting in two specialized medical VLMs: MIM-LLaVA-Med and Med-Mantis, both optimized for multi-image analysis. Additionally, we develop the Med-MIM benchmark to comprehensively evaluate the medical multi-image understanding capabilities of LVLMs. We assess eight popular LVLMs, including our two models, on the Med-MIM benchmark. Experimental results show that both Med-Mantis and MIM-LLaVA-Med achieve superior performance on the held-in and held-out subsets of the Med-MIM benchmark, demonstrating that the Med-MIM instruction dataset effectively enhances LVLMs' multi-image understanding capabilities in the medical domain.

医学视觉语言模型多图理解医疗问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。