arXiv:2602.23363cs.CV2026-02被引 6

用多信号奖励训练医疗大模型,让其能自由回答临床问题。

MediX-R1: Open Ended Medical Reinforcement Learning

  • 用分组强化学习与复合奖励优化医疗多模态模型
  • 仅用5.1万条指令数据,在开放问答任务上表现领先
  • 适合需要真实临床推理能力的医疗AI研究者

我们提出MediX-R1,一个面向医疗多模态大语言模型的开放式强化学习框架,支持超越选择题的临床实证式自由回答。该框架通过分组强化学习微调视觉-语言基础模型,并设计复合奖励机制:基于LLM的准确性奖励以严格是非判断语义正确性,基于医学嵌入的语义奖励捕捉同义表达与术语变体,以及轻量级格式与模态奖励,确保可解释推理与模态识别。这种多信号设计在传统可验证或仅限选择题奖励失效的开放输出场景中提供稳定、丰富的反馈。为评估进展,我们提出统一评估框架,采用基于参考的LLM作为裁判,替代脆弱的字符串匹配指标,有效衡量语义正确性、推理过程与上下文一致性。尽管仅使用约5.1万条指令样本,MediX-R1在标准医疗大模型(纯文本)和视觉语言模型(图像+文本)基准上均取得优异表现,显著超越主流开源基线,尤其在开放问答任务上提升明显。结果表明,结合全面奖励信号与LLM评估的开放式强化学习,是实现可靠医疗多模态推理的可行路径。训练模型、数据集及源代码已公开于https://medix.cvmbzuai.com。

原文摘要 · Abstract (English)

We introduce MediX-R1, an open-ended Reinforcement Learning (RL) framework for medical multimodal large language models (MLLMs) that enables clinically grounded, free-form answers beyond multiple-choice formats. MediX-R1 fine-tunes a baseline vision-language backbone with Group Based RL and a composite reward tailored for medical reasoning: an LLM-based accuracy reward that judges semantic correctness with a strict YES/NO decision, a medical embedding-based semantic reward to capture paraphrases and terminology variants, and lightweight format and modality rewards that enforce interpretable reasoning and modality recognition. This multi-signal design provides stable, informative feedback for open-ended outputs where traditional verifiable or MCQ-only rewards fall short. To measure progress, we propose a unified evaluation framework for both text-only and image+text tasks that uses a Reference-based LLM-as-judge in place of brittle string-overlap metrics, capturing semantic correctness, reasoning, and contextual alignment. Despite using only $\sim51$K instruction examples, MediX-R1 achieves excellent results across standard medical LLM (text-only) and VLM (image + text) benchmarks, outperforming strong open-source baselines and delivering particularly large gains on open-ended clinical tasks. Our results demonstrate that open-ended RL with comprehensive reward signals and LLM-based evaluation is a practical path toward reliable medical reasoning in multimodal models. Our trained models, curated datasets and source code are available at https://medix.cvmbzuai.com

医疗AI强化学习多模态开放问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。