arXiv:2505.23867cs.CLcs.AI2025-05被引 2

小数据也能练出强医学多模态模型,推理能力更强

InfiMed: Low-Resource Medical MLLMs with Advancing Understanding and Reasoning

  • 用通用+高质量文本数据增强医疗数据,提升基础理解能力
  • 合成反思式思维链,让模型具备结构化推理起点
  • 30亿参数模型在7个医学基准上领先大模型,适合医疗垂域应用

多模态大模型在视觉理解与数学推理方面进展显著,但在医疗领域受限于两大挑战:一是多模态医疗数据稀缺且信息稀疏,限制推理深度;二是可验证奖励的强化学习(RLVR)在通用领域有效,却难以提升医疗模型性能。为此,在监督微调(SFT)阶段,我们引入高质量文本推理数据与通用多模态数据,结合医疗数据,高效提升基础医学能力并恢复模型推理能力。针对数据稀疏问题,额外合成注入反思模式的思维链(CoT),赋予模型初始反思推理能力,为后续RLVR训练提供结构化基础。最终提出InfiMed系列模型,InfiMed-SFT-3B与InfiMed-RL-3B在七个多模态医疗基准上达到顶尖表现。其中,InfiMed-RL-3B平均准确率达59.2%,优于更大模型InternVL3-8B的57.3%。SFT阶段使用18.8万样本,RLVR阶段使用3.6万样本,验证了两种策略的有效性。大量实验提供了推动医疗多模态模型发展的关键洞见。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in domains such as visual understanding and mathematical reasoning. However, their application in the medical domain is constrained by two key challenges: (1) multimodal medical datasets are scarce and often contain sparse information, limiting reasoning depth; and (2) Reinforcement Learning with Verifiable Rewards (RLVR), though effective in general domains, cannot reliably improve model performance in the medical domain. To overcome these challenges, during the supervised fine-tuning (SFT) stage, we incorporate high-quality textual reasoning data and general multimodal data alongside multimodal medical data to efficiently enhance foundational medical capabilities and restore the base model's reasoning ability. Moreover, considering that there are some multimodal medical datasets with sparse information, we further synthesize reflective-pattern-injected chain-of-thought (CoT) in addition to general CoT samples, equipping the model with initial reflective reasoning capabilities that provide a structured foundation for subsequent RLVR training. Finally, we introduce our InfiMed-Series models, InfiMed-SFT-3B and InfiMed-RL-3B, both of which deliver state-of-the-art performance across seven multimodal medical benchmarks. Notably, InfiMed-RL-3B achieves an average accuracy of 59.2%, outperforming even larger models like InternVL3-8B, which achieves 57.3%. Specifically, during the SFT phase, we utilized 188K samples, while the RLVR phase incorporated 36K samples, demonstrating the efficacy of both training strategies in achieving superior performance. We also conducted a series of extensive experiments, which provide valuable insights that contribute to advancing the performance of MLLMs in medical scenarios.

医学多模态小样本训练思维链强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。