解决医学开放问答中奖励坍塌问题,提升模型推理可靠性。
Adaptive Reinforcement for Open-ended Medical Reasoning via Semantic-Guided Reward Collapse Mitigation
- 用思维链监督微调注入医学知识,再结合自适应语义奖励优化
- 在6个医疗问答数据集上准确率显著提升,泛化能力更强
- 适合构建临床级多模态推理系统,尤其关注真实诊断流程
基于规则的强化学习(RL)在提升视觉语言模型(VLMs)推理深度和泛化能力方面展现出巨大潜力,同时保持计算效率。尽管如此,其在医学影像领域的应用仍有限。现有研究主要集中在封闭式视觉问答(VQA),难以适配真实的临床推理场景。而开放式医学VQA更贴近实际诊断流程,但研究较少。尽管已有研究尝试通过语义引导的RL实现格式衔接,但模型驱动的语义奖励常出现奖励坍塌,即不同语义的回答得分趋同。为此,本文提出针对开放式医学VQA的自适应强化框架ARMed。该方法先在思维链标注数据上进行监督微调(SFT),再通过文本正确性与自适应语义奖励进行强化优化,提升推理一致性和事实准确性。在六个挑战性医疗VQA基准上的实验表明,ARMed显著提升了准确率与泛化能力。结果强调了奖励可区分性在医学强化学习中的重要性,也展示了自适应语义奖励在构建鲁棒、临床可信的多模态推理系统中的潜力。
原文摘要 · Abstract (English)
Reinforcement learning (RL) with rule-based reward functions has recently shown great promise in enhancing the reasoning depth and generalization ability of vision-language models (VLMs), while maintaining computational efficiency. In spite of these advances, its adoption in medical imaging remains limited. Current reinforcement fine-tuning (RFT) efforts in this field mainly focus on closed-ended visual question answering (VQA), restricting their applicability to realistic clinical reasoning. However, open-ended medical VQA better mirrors clinical diagnostic workflows but remains underexplored. Although several studies have attempted to bridge the two formats through semantically guided RL, model-driven semantic rewards often suffer from reward collapse, where responses with distinct semantics yield nearly identical scores. To overcome this limitation, we introduce Adaptive Reinforcement for Medical Reasoning (ARMed), a novel RL framework tailored for open-ended medical VQA. ARMed first injects domain expertise through supervised fine-tuning (SFT) on chain-of-thought annotations, followed by reinforcement optimization using textual correctness and adaptive semantic rewards to refine reasoning consistency and factual accuracy. Extensive experiments on six challenging medical VQA benchmarks demonstrate that ARMed substantially improves both accuracy and generalization. These findings underscore the importance of reward discriminability in medical RL and highlight the potential of adaptive semantic rewards for building robust, clinically reliable multimodal reasoning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。