用强化学习提升医疗推理能力,让AI像专家一样思考并解释决策。
Fleming-R1: Toward Expert-Level Medical Reasoning via Reinforcement Learning
- 构建医疗知识图引导的数据策略,增强罕见病和复杂推理覆盖。
- 通过思维链冷启动与可验证奖励强化学习,显著提升推理准确性。
- 小模型表现超越大模型,适合临床安全部署的高可信AI系统。
尽管大语言模型在医疗应用中展现潜力,但实现专家级临床推理仍面临挑战,需兼顾答案准确性和推理过程透明性。为此,我们提出Fleming-R1,通过三项互补创新实现可验证的医疗推理:首先,采用面向推理的数据策略(RODS),结合精选医疗问答数据集与知识图谱引导生成,提升罕见疾病、药物及多跳推理链的覆盖率;其次,利用思维链冷启动从教师模型中蒸馏高质量推理路径,建立稳健的推理先验;第三,设计两阶段基于可验证奖励的强化学习框架(RLVR),采用组相对策略优化,并通过自适应难例挖掘解决顽固错误模式。在多个医疗基准测试中,Fleming-R1表现出显著的参数高效改进:7B版本超越更大规模基线,32B模型接近GPT-4o水平,且持续优于主流开源模型。结果表明,结构化数据设计、推理导向初始化与可验证强化学习能推动临床推理超越单纯准确率优化。我们公开发布Fleming-R1,以促进医疗AI的透明、可复现与可审计发展,支持高风险临床环境中的安全应用。
原文摘要 · Abstract (English)
While large language models show promise in medical applications, achieving expert-level clinical reasoning remains challenging due to the need for both accurate answers and transparent reasoning processes. To address this challenge, we introduce Fleming-R1, a model designed for verifiable medical reasoning through three complementary innovations. First, our Reasoning-Oriented Data Strategy (RODS) combines curated medical QA datasets with knowledge-graph-guided synthesis to improve coverage of underrepresented diseases, drugs, and multi-hop reasoning chains. Second, we employ Chain-of-Thought (CoT) cold start to distill high-quality reasoning trajectories from teacher models, establishing robust inference priors. Third, we implement a two-stage Reinforcement Learning from Verifiable Rewards (RLVR) framework using Group Relative Policy Optimization, which consolidates core reasoning skills while targeting persistent failure modes through adaptive hard-sample mining. Across diverse medical benchmarks, Fleming-R1 delivers substantial parameter-efficient improvements: the 7B variant surpasses much larger baselines, while the 32B model achieves near-parity with GPT-4o and consistently outperforms strong open-source alternatives. These results demonstrate that structured data design, reasoning-oriented initialization, and verifiable reinforcement learning can advance clinical reasoning beyond simple accuracy optimization. We release Fleming-R1 publicly to promote transparent, reproducible, and auditable progress in medical AI, enabling safer deployment in high-stakes clinical environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。