arXiv:2505.23118cs.CLcs.AI2025-05被引 7

用两阶段训练提升医疗多模态推理能力,效果显著优于基线。

Elicit and Enhance: Advancing Multimodal Reasoning in Medical Scenarios

  • 分两阶段训练:先用2000条文本示范激发推理行为,再用1500个医学案例优化多模态推理偏好。
  • 在多个医疗多模态基准上,训练后的模型性能持续超越基线,大模型和推理时扩展也表现稳健。
  • 适合医学AI研究者和临床辅助系统开发者,推动多模态医疗决策智能化。

有效的临床决策依赖于对多样化证据的迭代式多模态推理。尽管多模态推理模型在数学与科学任务中取得显著进展,其在医疗领域的应用仍相对不足。本文提出MedE²,一种两阶段后训练流程,用于激发并增强医疗场景下的多模态推理能力。第一阶段通过2,000条精心设计的纯文本推理示范微调模型,以激发推理行为;第二阶段利用1,500个严格筛选的多模态医学案例,对齐模型输出与提出的多模态医疗推理偏好。大量实验表明,MedE²能有效提升医疗多模态模型的推理性能。值得注意的是,经由MedE²训练的模型在多个医疗多模态基准上均持续优于基线模型,且在更大规模模型及推理时缩放条件下仍具鲁棒性与实用性。

原文摘要 · Abstract (English)

Effective clinical decision-making depends on iterative, multimodal reasoning across diverse sources of evidence. The recent emergence of multimodal reasoning models has significantly transformed the landscape of solving complex tasks. Although such models have achieved notable success in mathematics and science, their application to medical domains remains underexplored. In this work, we propose \textit{MedE$^2$}, a two-stage post-training pipeline that elicits and then enhances multimodal reasoning for medical domains. In Stage-I, we fine-tune models using 2,000 text-only data samples containing precisely orchestrated reasoning demonstrations to elicit reasoning behaviors. In Stage-II, we further enhance the model's reasoning capabilities using 1,500 rigorously curated multimodal medical cases, aligning model reasoning outputs with our proposed multimodal medical reasoning preference. Extensive experiments demonstrate the efficacy and reliability of \textit{MedE$^2$} in improving the reasoning performance of medical multimodal models. Notably, models trained with \textit{MedE$^2$} consistently outperform baselines across multiple medical multimodal benchmarks. Additional validation on larger models and under inference-time scaling further confirms the robustness and practical utility of our approach.

医疗AI多模态推理模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。