用可验证的临床结局评分,让大模型更像医生一样做住院决策。
CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics
- 将住院决策建模为部分可观测马尔可夫决策过程,用未来可验证的评分指导模型。
- 在自建数据集上达84.91%准确率,优于GPT-5和MedGemma-27B等前沿模型。
- 经600+名医生验证,适合医疗大模型临床推理能力评估与部署。
住院临床推理是部分可观测下的序列决策:医生基于当前入院信息需选择下一步行动,而其远期后果尚未显现。现有临床大模型评估与强化学习奖励信号存在闭合检索、诊疗路径泄露或无锚定的模型自评问题。本文提出CLR-voyance框架,将住院推理重构为部分可观测马尔可夫决策过程(POMDP),并引入同时基于临床结局与医生验证的奖励机制。具体实现为CLR-POMDP,将成功病程划分为策略可见过去与仅由全知模型知晓的未来。基于过去信息,全知大模型生成病例特异的问答对,形成首个可在未来验证的临床推理评判标准。该标准用于模型后训练与评估。我们对Qwen3-8B和MedGemma-4B采用GRPO后模型融合进行后训练,在保留通用能力的同时达成业界最优的住院临床推理性能。CLR-voyance-8B在CLR-POMDP上达到84.91%准确率,超越GPT-5(77.83%)和MedGemma-27B(66.66%),且在现有医学基准上表现相当或更优。为确保临床意义,我们开展大规模医生对齐研究,邀请医师定制每例评判标准、评分候选回复,并盲评模型推理偏好。研究揭示了临床大模型作为评判者与偏好模型选择的关键洞见。该系统已在合作公立医院部署超6个月,协助撰写数千份高复杂度住院记录。
原文摘要 · Abstract (English)
Inpatient clinical reasoning is a sequential decision under partial observability: the clinician sees the admission so far and must choose the next action whose downstream consequences are not yet visible. Existing clinical-LLM evaluations and RL rewards signals collapse this into closed-form retrieval, clinical journey leakage, or unanchored LLM-as-judge scoring. We introduce CLR-voyance, a framework that reformulates inpatient reasoning as a Partially Observable Markov Decision Process (POMDP) and supervises it with rewards that are simultaneously outcome-grounded and clinician-validated. We instantiate the formulation as CLR-POMDP, which partitions successful patient journeys into a policy-visible past and an oracle-only future. Using the past information, an oracle LLM generates a case-specific query-answer pair, and the first adaptive rubric for clinical reasoning which is verifiable in the future of the patient journey. These rubrics are used for both post-training and evaluation of models for inpatient clinical reasoning. We post-train Qwen3-8B and MedGemma-4B with GRPO followed by model merging, yielding state-of-the-art inpatient clinical reasoning while retaining generalist capabilities. CLR-voyance-8B achieves 84.91% on CLR-POMDP, ahead of frontier medical reasoning models like GPT-5 (77.83%) and MedGemma-27B (66.66%) and has comparable or better performance on existing medical benchmarks. To ensure a clinically meaningful setting, we conduct a large-scale clinician alignment study, where physicians curate per-case rubrics, grade candidate responses, and provide blinded pairwise preferences of model reasoning. This study provides insights on clinical LLM-as-a-judge and clinical preference-model selection, which can inform the community at large. CLR-voyance has been deployed for 6+ months at a partner public hospital, drafting thousands of reasoning-heavy inpatient notes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。