不依赖标注推理数据,用简单规则奖励让医学大模型自主学会推理。
Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL
- 用公开多选题数据+极简规则奖励训练,无需人工标注推理链。
- 在6个医学问答基准上达到顶尖水平,超越更大或闭源模型。
- 揭示数据质量比数量更重要,适合医疗AI可解释性研究者参考。
提升大语言模型在复杂任务中的表现并实现可解释决策,尤其在临床应用中,需有效推理能力。然而,在缺乏昂贵链式思维(CoT)数据监督微调的情况下,这仍具挑战性。本文提出AlphaMed,首个仅通过强化学习在公开多选题数据集上实现推理能力的医学大模型,无需监督微调或从闭源模型(如GPT-4o)蒸馏的CoT数据。AlphaMed在六个医学问答基准上取得领先性能,优于传统SFT+RL训练模型。在困难基准(如MedXpert)上,甚至超过更大或闭源模型(如DeepSeek-V3-671B和Claude-3.5-Sonnet)。我们通过三个问题展开数据驱动分析:(i) 极简规则奖励能否在无蒸馏CoT监督下激励推理?(ii) 数据量与多样性如何影响推理?(iii) 问题难度如何塑造推理的出现与泛化?结果表明,数据信息量是推理性能的关键驱动力,且在信息丰富、多选题数据上的极简强化学习能有效诱导推理,无需CoT监督。不同基准表现出异质趋势,凸显当前评估局限,亟需更难、更注重推理的医学问答基准。
原文摘要 · Abstract (English)
Improving performance on complex tasks and enabling interpretable decision making in large language models (LLMs), especially for clinical applications, requires effective reasoning. Yet this remains challenging without supervised fine-tuning (SFT) on costly chain-of-thought (CoT) data distilled from closed-source models (e.g., GPT-4o). In this work, we present AlphaMed, the first medical LLM to show that reasoning capability can emerge purely through reinforcement learning (RL), using minimalist rule-based rewards on public multiple-choice QA datasets, without relying on SFT or distilled CoT data. AlphaMed achieves state-of-the-art results on six medical QA benchmarks, outperforming models trained with conventional SFT+RL pipelines. On challenging benchmarks (e.g., MedXpert), AlphaMed even surpasses larger or closed-source models such as DeepSeek-V3-671B and Claude-3.5-Sonnet. To understand the factors behind this success, we conduct a comprehensive data-centric analysis guided by three questions: (i) Can minimalist rule-based RL incentivize reasoning without distilled CoT supervision? (ii) How do dataset quantity and diversity impact reasoning? (iii) How does question difficulty shape the emergence and generalization of reasoning? Our findings show that dataset informativeness is a key driver of reasoning performance, and that minimalist RL on informative, multiple-choice QA data is effective at inducing reasoning without CoT supervision. We also observe divergent trends across benchmarks, underscoring limitations in current evaluation and the need for more challenging, reasoning-oriented medical QA benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。