arXiv:2505.19213cs.AI2025-05被引 18

用分阶段强化学习提升医疗视觉问答的推理能力。

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning

  • 分阶段训练:先学封闭式问题,再过渡到开放式推理。
  • 在8个基准上平均提升5.7%准确率,域内最高增11.4%。
  • 适合需要临床可解释性推理的医疗AI研究者。

近期基于可验证规则奖励的强化学习显著提升了视觉语言模型的推理能力与分布外泛化性能,无需人工设计推理链。尽管通用领域进展显著,其在医学影像中的应用仍受限。现有医学强化微调方法多聚焦封闭式VQA,限制了模型获取世界知识和灵活适应任务的能力,更难以满足临床所需的开放式、高推理强度决策需求。为此,我们提出首个面向医学VQA的多模态强化学习框架MedCCO,通过课程驱动的微调范式统一封闭式与开放式数据。MedCCO首先在多样化的封闭式医学VQA任务上进行微调,建立领域基础推理能力;随后逐步迁移到开放式任务,促进深层知识增强与临床可解释性。我们在八个具有挑战性的医学VQA基准上验证了MedCCO,涵盖封闭式与开放式场景。实验结果表明,MedCCO持续提升性能与泛化能力,在三个域内任务中平均提升11.4%准确率,在五个域外基准上提升5.7%。这些发现凸显了课程引导强化学习在推动鲁棒、临床相关推理方面的潜力。

原文摘要 · Abstract (English)

Recent advances in reinforcement learning with verifiable, rule-based rewards have greatly enhanced the reasoning capabilities and out-of-distribution generalization of VLMs/LLMs, obviating the need for manually crafted reasoning chains. Despite these promising developments in the general domain, their translation to medical imaging remains limited. Current medical reinforcement fine-tuning (RFT) methods predominantly focus on close-ended VQA, thereby restricting the model's ability to engage in world knowledge retrieval and flexible task adaptation. More critically, these methods fall short of addressing the critical clinical demand for open-ended, reasoning-intensive decision-making. To bridge this gap, we introduce \textbf{MedCCO}, the first multimodal reinforcement learning framework tailored for medical VQA that unifies close-ended and open-ended data within a curriculum-driven RFT paradigm. Specifically, MedCCO is initially fine-tuned on a diverse set of close-ended medical VQA tasks to establish domain-grounded reasoning capabilities, and is then progressively adapted to open-ended tasks to foster deeper knowledge enhancement and clinical interpretability. We validate MedCCO across eight challenging medical VQA benchmarks, spanning both close-ended and open-ended settings. Experimental results show that MedCCO consistently enhances performance and generalization, achieving a 11.4\% accuracy gain across three in-domain tasks, and a 5.7\% improvement on five out-of-domain benchmarks. These findings highlight the promise of curriculum-guided RL in advancing robust, clinically-relevant reasoning in medical multimodal language models.

医疗推理强化学习多模态VQA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。