arXiv:2511.22232cs.CVcs.AI2025-11被引 1

用医学文献中的复合图像训练多图理解模型,提升临床诊断能力

From Compound Figures to Composite Understanding: Developing a Multi-Modal LLM from Biomedical Literature with Medical Multiple-Image Benchmarking and Validation

  • 通过分解复合图像任务,让模型学习跨模态、时空关系
  • 基于23.7万张复合图构建M3LLM,在多图任务上超越现有模型
  • 适合医疗AI研究者与临床辅助系统开发者使用

多模态大语言模型在医疗领域展现潜力,但多数模型仍局限于单图理解,难以支持临床中需整合多模态或时序图像的诊断需求。现有模型受限于缺乏大规模高质量标注数据。为此,我们提出新框架,利用生物医学文献中可自由使用的复合图像作为丰富但未被充分利用的数据源。设计五阶段上下文感知指令生成范式,采用分而治之策略,将多图分析拆解为可管理子任务,使模型能学习复杂的空间、时间与跨模态关系,实现综合理解。通过对超过23.7万张复合图像及其上下文文本进行解析,构建M3LLM——一个医学多图多模态大语言模型。为评估性能,我们建立人工专家验证的PMC-MI-Bench基准。实验表明,M3LLM在多图、单图、纯文本及多选任务中均显著优于通用和专用医疗多模态模型。尤其在使用MIMIC数据集进行纵向胸片分析时表现出强泛化能力。本工作建立了一种可扩展、高效的医学多模态模型开发范式,弥合了生物医学文献与临床应用之间的鸿沟。

原文摘要 · Abstract (English)

Multi-modal large language models (MLLMs) have shown promise in advancing healthcare. However, most existing models remain confined to single-image understanding, which greatly limits their applicability in clinical workflows. In practice, medical diagnosis and progression often require synthesizing information across multiple images from different modalities or time points. The development of medical MLLMs capable of such multi-image understanding has been hindered by the lack of large-scale, high-quality annotated training data. To address this limitation, we propose a novel framework that leverages license-permissive compound images in biomedical literature, as a rich yet underutilized data source for multi-image analysis. Specifically, we design a five-stage, context-aware instruction generation paradigm underpinned by a divide-and-conquer strategy. By decomposing multi-image analysis into manageable sub-tasks, this paradigm empowers MLLMs to move beyond single-panel analysis and provide a composite understanding by learning the complex spatial, temporal, and cross-modal relationships inherent in these compound figures. By parsing over 237,000 compound figures and their contextual text for instruction generation, we develop M3LLM, a medical multi-image multi-modal large language model. For benchmarking, we construct PMC-MI-Bench for composite understanding, manually validated by medical experts. Extensive experiments show that M3LLM significantly outperforms both general-purpose and specialized medical MLLMs across multi-image, single-image, text-only, and multi-choice scenarios. Notably, M3LLM exhibits strong generalization to longitudinal chest X-ray analysis using the MIMIC dataset. This work establishes a scalable and efficient paradigm for developing medical MLLMs capable of composite reasoning, bridging the gap between biomedical literature and real-world clinical applications.

多模态模型医学图像复合理解AI医疗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。