arXiv:2602.06965cs.CV2026-02被引 10

MedMO提升医学图像多模态模型的对齐与推理能力,性能超越现有开源模型。

MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images

  • 采用分阶段训练,融合跨模态预训练与强化学习优化空间定位和逻辑推理
  • 在多个医学任务上显著领先,如VQA提升21.3%,报告生成提升6.7%
  • 适用于放射、眼科、病理等多种医学影像场景,适合临床研究与AI辅助诊断

多模态大语言模型发展迅速,但在医学领域受限于领域覆盖不足、模态对齐不完善及缺乏可解释推理。我们提出MedMO,一种基于通用多模态大模型架构、仅在大规模医学数据上训练的医疗多模态基础模型。MedMO采用多阶段训练:首先通过跨模态预训练对齐异构视觉编码器与医学语言主干;其次使用多任务指令微调,涵盖图像描述、视觉问答、报告生成、检索与病灶定位;最后引入强化学习,结合事实性验证与框级GIoU信号,提升复杂临床场景下的空间对齐与逐步推理能力。在多种模态与任务上,MedMO均超越强基线模型。MedMO-8B-Next在视觉问答基准上平均提升6.6%,其中在MMMU-Med提升6.0%、PMC-VQA提升9.8%、MedXpertQA提升21.3%;文本问答提升14.4%,其中MMLU-Med提升8.4%、MedQA提升30.1%;报告生成提升6.7%(MIMIC-CXR)。在定位任务中,其在Bacteria数据集达到56.1 IoU,较Fleming-VL-8B提升47.8。小规模版本MedMO-4B-Next仍具竞争力,全面优于Fleming-VL-8B。跨放射、眼科、病理显微镜等模态评估进一步验证其泛化能力。项目地址:https://genmilab.github.io/MedMO-Page

原文摘要 · Abstract (English)

Multimodal large language models have advanced rapidly, but their adoption in medicine is constrained by limited domain coverage, imperfect modality alignment, and insufficient grounded reasoning. We introduce MedMO, a medical multimodal foundation model built on a general MLLM architecture and trained exclusively on large-scale domain-specific data. MedMO uses a multi-stage training recipe that includes cross-modal pretraining to align heterogeneous visual encoders with a medical language backbone, instruction tuning with multi-task supervision spanning captioning, VQA, report generation, retrieval, and bounding-box disease localization, and reinforcement learning with verifiable rewards that combine factuality checks with a box-level GIoU signal to improve spatial grounding and step-by-step reasoning in challenging clinical settings. Across modalities and tasks, MedMO surpasses strong open-source medical baselines. MedMO-8B-Next achieves consistent gains on VQA benchmarks, improving by 6.6% on average over Fleming-VL-8B, including gains of 6.0% on MMMU-Med, 9.8% on PMC-VQA, and 21.3% on MedXpertQA. On text-based QA, it improves by 14.4% over Fleming-VL-8B, driven by gains of 8.4% on MMLU-Med and 30.1% on MedQA. For medical report generation, it improves by 6.7% on MIMIC-CXR. MedMO-8B-Next also demonstrates strong grounding performance, reaching 56.1 IoU on Bacteria, which is a 47.8 IoU gain over Fleming-VL-8B. At smaller scale, MedMO-4B-Next remains competitive and exceeds Fleming-VL-8B across VQA, QA, and report generation. Evaluations spanning radiology, ophthalmology, and pathology microscopy further confirm broad cross-modality generalization. Project is available at https://genmilab.github.io/MedMO-Page

医学多模态视觉问答空间对齐大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。