arXiv:2503.13939cs.CV2025-03被引 150

用强化学习提升医疗视觉语言模型的推理能力与泛化性。

Med-R1: Reinforcement Learning for Generalizable Medical Reasoning in Vision-Language Models

论文配图:Med-R1: Reinforcement Learning for Generalizable Medical Reasoning in Vision-Language Models
图 1 · 摘自论文原文
  • 基于强化学习优化医疗视觉语言模型,避免依赖大量标注数据。
  • 在8种医学影像任务中准确率比基线高29.94%,超越36倍参数的大模型。
  • 发现高质量推理比多步骤思考更重要,适合临床部署的高效模型设计。

视觉语言模型(VLMs)在自然图像推理中取得显著进展,但在医学影像领域的潜力尚未充分挖掘。医学视觉语言任务要求精确理解与临床一致的答案,但受限于医学数据复杂性和高质量专家标注稀缺,传统监督微调(SFT)和思维链(CoT)策略效果有限。为此,我们提出Med-R1,一种基于强化学习(RL)的视觉语言模型,旨在提升医学推理的泛化性与可靠性。基于DeepSeek策略,Med-R1采用组相对策略优化(GRPO),推动奖励驱动的学习超越静态标注。我们在八种不同医学影像模态上全面评估Med-R1,其平均准确率相比基线模型Qwen2-VL-2B提升29.94%,甚至优于参数量大36倍的Qwen2-VL-72B-a模型。为评估跨任务泛化能力,我们在五类问题上测试,其在问题类型泛化上比Qwen2-VL-2B高出32.06%,也超过Qwen2-VL-72B。进一步分析发现,去除中间推理过程(No-Thinking-Med-R1)不仅提升域内与跨域泛化,且训练更少,挑战了‘更多推理一定更好’的假设。结果表明,在医学VQA中,决定有效性的不是推理本身,而是其质量与领域一致性。综合来看,强化学习能有效提升医学推理与泛化能力,推动高效可靠的视觉语言模型在真实场景中的应用。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have achieved impressive progress in natural image reasoning, yet their potential in medical imaging remains underexplored. Medical vision-language tasks demand precise understanding and clinically coherent answers, which are difficult to achieve due to the complexity of medical data and the scarcity of high-quality expert annotations. These challenges limit the effectiveness of conventional supervised fine-tuning (SFT) and Chain-of-Thought (CoT) strategies that work well in general domains. To address these challenges, we propose Med-R1, a reinforcement learning (RL)-enhanced vision-language model designed to improve generalization and reliability in medical reasoning. Built on the DeepSeek strategy, Med-R1 adopts Group Relative Policy Optimization (GRPO) to encourage reward-guided learning beyond static annotations. We comprehensively evaluate Med-R1 across eight distinct medical imaging modalities. Med-R1 achieves a 29.94% improvement in average accuracy over its base model Qwen2-VL-2B, and even outperforms Qwen2-VL-72B-a model with 36x more parameters. To assess cross-task generalization, we further evaluate Med-R1 on five question types. Med-R1 outperforms Qwen2-VL-2B by 32.06% in question-type generalization, also surpassing Qwen2-VL-72B. We further explore the thinking process in Med-R1, a crucial component for the success of Deepseek-R1. Our results show that omitting intermediate rationales (No-Thinking-Med-R1) not only improves in-domain and cross-domain generalization with less training, but also challenges the assumption that more reasoning always helps. These findings suggest that in medical VQA, it is not reasoning itself, but its quality and domain alignment, that determine effectiveness. Together, these results highlight that RL improves medical reasoning and generalization, enabling efficient and reliable VLMs for real-world deployment.

医疗视觉语言强化学习模型泛化医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。