arXiv:2509.02208cs.LGcs.AI2025-09被引 52

用动态验证系统提升医疗大模型临床实用能力

Baichuan-M2: Scaling Medical Capability with Large Verifier System

  • 构建患者模拟器与动态评分生成器,实现真实诊疗交互环境
  • 320亿参数模型在HealthBench Hard上得分超32,逼近GPT-5水平
  • 适合关注医疗AI落地、临床决策支持的研究者与开发者

随着大语言模型在对话与推理能力上的进步,其在医疗领域的实际应用成为研究重点。然而,现有医疗LLM在静态考试(如USMLE)上的表现与其在真实临床决策中的实用性之间存在显著差距,原因在于传统评测无法反映诊疗过程的动态交互特性。为此,本文提出一种新型动态验证框架,超越静态答案校验,构建大规模高保真交互式强化学习系统。该框架包含两个核心组件:基于去标识化病历数据的患者模拟器,以及可动态生成多维评估指标的临床评分生成器。在此基础上,我们开发了320亿参数的医疗增强推理模型Baichuan-M2,采用改进的分组相对策略优化(GRPO)算法进行多阶段强化学习训练。在HealthBench评测中,Baichuan-M2超越所有开源模型及多数闭源先进模型,在HealthBench Hard基准上得分超过32,此前仅有GPT-5达到此水平。研究表明,强大的动态验证系统对于对齐大模型能力与临床实际应用至关重要,为医疗AI部署建立了新的性能-参数权衡帕累托前沿。

原文摘要 · Abstract (English)

As large language models (LLMs) advance in conversational and reasoning capabilities, their practical application in healthcare has become a critical research focus. However, there is a notable gap between the performance of medical LLMs on static benchmarks such as USMLE and their utility in real-world clinical decision-making. This discrepancy arises because traditional exams fail to capture the dynamic, interactive nature of medical consultations. To address this challenge, we introduce a novel dynamic verification framework that moves beyond static answer verifier, establishing a large-scale, high-fidelity interactive reinforcement learning system. Our framework comprises two key components: a Patient Simulator that creates realistic clinical environments using de-identified medical records, and a Clinical Rubrics Generator that dynamically produces multi-dimensional evaluation metrics. Building on this foundation, we develop Baichuan-M2, a 32B-parameter medical augmented reasoning model trained through a multi-stage reinforcement learning strategy with an improved Group Relative Policy Optimization (GRPO) algorithm. Evaluated on HealthBench, Baichuan-M2 outperforms all other open-source models and most advanced closed-source counterparts, achieving a score above 32 on the challenging HealthBench Hard benchmark-previously exceeded only by GPT-5. Our work demonstrates that robust dynamic verifier system is essential for aligning LLM capabilities with practical clinical applications, establishing a new Pareto front in the performance-parameter trade-off for medical AI deployment.

医疗AI大模型强化学习临床决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。