arXiv:2505.13973cs.CLcs.AI2025-05被引 8

用强化学习提升医疗视觉问答模型的临床推理能力

Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models

  • 基于GRPO的强化学习优化医疗视觉问答
  • 相比监督微调,准确率与推理质量双提升
  • 揭示医学对齐、长度奖励等关键影响因素

近期,基于强化学习(RL)的调优改变了多模态大语言模型(MLLMs)的发展方向,尤其是组相对策略优化(GRPO)的提出。然而,直接将该方法应用于医疗任务仍面临挑战,难以实现符合临床预期的模型行为。为使模型响应更贴近临床需求,我们研究了影响医疗视觉问答(VQA)中基于强化学习调优有效性的四个关键维度:基础模型初始化策略、医学语义对齐的作用、基于长度的奖励对长链推理的影响,以及偏见的影响。通过大量实验分析这些因素在医疗MLLMs中的作用,提供了领域特定微调的新见解。结果表明,基于GRPO的强化学习调优在准确性和推理质量上均显著优于标准监督微调(SFT)。

原文摘要 · Abstract (English)

Recently, reinforcement learning (RL)-based tuning has shifted the trajectory of Multimodal Large Language Models (MLLMs), particularly following the introduction of Group Relative Policy Optimization (GRPO). However, directly applying it to medical tasks remains challenging for achieving clinically grounded model behavior. Motivated by the need to align model response with clinical expectations, we investigate four critical dimensions that affect the effectiveness of RL-based tuning in medical visual question answering (VQA): base model initialization strategy, the role of medical semantic alignment, the impact of length-based rewards on long-chain reasoning, and the influence of bias. We conduct extensive experiments to analyze these factors for medical MLLMs, providing new insights into how models are domain-specifically fine-tuned. Additionally, our results also demonstrate that GRPO-based RL tuning consistently outperforms standard supervised fine-tuning (SFT) in both accuracy and reasoning quality.

医疗VQA强化学习多模态模型临床对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。