arXiv:2505.17387cs.CL2025-05被引 1

320亿参数医学大模型用强化学习提升诊断推理能力

WiNGPT-3.0 Technical Report

  • 分阶段训练融合监督微调与强化学习,基于链式思维数据提升推理能力
  • 在MedQA-USMLE上达87.1分,在临床推理任务中从58.1提升至62.5
  • 仅用数千样本强化学习即有效,适合医疗系统部署和可信推理

当前大语言模型在结构化、可解释、可验证的医学推理方面存在显著局限,且面临计算资源与数据隐私的部署挑战。本报告聚焦于320亿参数的WiNGPT-3.0模型开发,旨在增强其医学推理能力,并探索其在医疗IT基础设施中的集成潜力,推动临床可用模型的发展。方法采用多阶段训练流程,涵盖通用、医学与临床推理任务,结合监督微调(SFT)与强化学习(RL),利用精心构建的长链式思维(Long Chain-of-Thought, CoT)数据集、辅助奖励模型及基于证据的诊断链模拟。WiNGPT-3.0表现优异:特定版本在MedCalc上得分66.6,在MedQA-USMLE上达87.1;针对临床推理任务的专项训练使其得分从基线58.1提升至62.5。结果表明,即使仅使用数千样本的有限数据集,强化学习仍能有效提升医学推理准确性。这一发现证明了在数据与算力受限条件下,强化学习仍具可行性,为临床工作流与健康信息体系中更可信、可部署的大模型铺平道路。

原文摘要 · Abstract (English)

Current Large Language Models (LLMs) exhibit significant limitations, notably in structured, interpretable, and verifiable medical reasoning, alongside practical deployment challenges related to computational resources and data privacy. This report focused on the development of WiNGPT-3.0, the 32-billion parameter LLMs, engineered with the objective of enhancing its capacity for medical reasoning and exploring its potential for effective integration within healthcare IT infrastructures. The broader aim is to advance towards clinically applicable models. The approach involved a multi-stage training pipeline tailored for general, medical, and clinical reasoning. This pipeline incorporated supervised fine-tuning (SFT) and reinforcement learning (RL), leveraging curated Long Chain-of-Thought (CoT) datasets, auxiliary reward models, and an evidence-based diagnostic chain simulation. WiNGPT-3.0 demonstrated strong performance: specific model variants achieved scores of 66.6 on MedCalc and 87.1 on MedQA-USMLE. Furthermore, targeted training improved performance on a clinical reasoning task from a baseline score of 58.1 to 62.5. These findings suggest that reinforcement learning, even when applied with a limited dataset of only a few thousand examples, can enhance medical reasoning accuracy. Crucially, this demonstration of RL's efficacy with limited data and computation paves the way for more trustworthy and practically deployable LLMs within clinical workflows and health information infrastructures.

医学大模型强化学习临床推理可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。