arXiv:2502.19655cs.CLcs.AI2025-02被引 38

用强化学习让小模型自动生成医学推理,效果不输传统方法。

Med-RLVR: Emerging Medical Reasoning from a 3B base model via reinforcement Learning

  • 通过可验证奖励机制,让30亿参数模型在无推理标注下自主发展医学推理能力。
  • 在医学多选题上表现接近监督微调,分布外泛化能力提升8个百分点。
  • 首次证明强化学习可激发医学领域推理,适合医疗AI研究者参考。

基于可验证奖励的强化学习(RLVR)近期受到关注,因其能在无需显式推理监督的情况下,从基础语言模型中激发自我演化的推理能力,如DeepSeek-R1所示。尽管先前的RLVR研究主要集中在数学和编程领域,其在其他任务和领域的适用性仍未知。本文首次探索了医学推理是否可通过RLVR实现。我们提出了Med-RLVR,利用医学多选题问答(MCQA)数据作为可验证标签,在30亿参数基模型上开展研究。结果表明,RLVR不仅适用于数学和编程,也能成功应用于医学问答任务。值得注意的是,Med-RLVR在分布内任务上的性能与传统监督微调(SFT)相当,且在分布外泛化上显著提升,准确率提高8个百分点。进一步训练动态分析显示,即使没有显式推理监督,推理能力仍从3B参数基模型中自然涌现。这些发现突显了RLVR在数学和编程之外领域的潜力,为知识密集型领域(如医学)开辟了新路径。

原文摘要 · Abstract (English)

Reinforcement learning from verifiable rewards (RLVR) has recently gained attention for its ability to elicit self-evolved reasoning capabilitie from base language models without explicit reasoning supervisions, as demonstrated by DeepSeek-R1. While prior work on RLVR has primarily focused on mathematical and coding domains, its applicability to other tasks and domains remains unexplored. In this work, we investigate whether medical reasoning can emerge from RLVR. We introduce Med-RLVR as an initial study of RLVR in the medical domain leveraging medical multiple-choice question answering (MCQA) data as verifiable labels. Our results demonstrate that RLVR is not only effective for math and coding but also extends successfully to medical question answering. Notably, Med-RLVR achieves performance comparable to traditional supervised fine-tuning (SFT) on in-distribution tasks while significantly improving out-of-distribution generalization, with an 8-point accuracy gain. Further analysis of training dynamics reveals that, with no explicit reasoning supervision, reasoning emerges from the 3B-parameter base model. These findings underscore the potential of RLVR in domains beyond math and coding, opening new avenues for its application in knowledge-intensive fields such as medicine.

医学推理强化学习小模型自进化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。