arXiv:2505.24105cs.CL2025-05被引 12

用强化学习提升大模型在病历推理中的准确性和可解释性。

Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning

  • 先用轻量微调注入缺失医学知识,再用可验证奖励强化学习优化决策。
  • 在多个临床任务上实现准确率提升,且输出更结构化可读。
  • 适合想提升医疗大模型推理能力的研究者和开发者。

我们提出EHRMIND,一种基于可验证奖励的强化学习(RLVR)方法,用于将大语言模型(LLMs)适配至复杂的临床推理任务。尽管RLVR在数学与编程中表现成功,但在医疗领域因专业性强而面临挑战。初步实验显示两大失败模式:(1)知识误用,即模型拥有相关医学知识但应用错误;(2)知识缺失,即缺少关键领域知识。EHRMIND采用两阶段策略:首先进行轻量级监督微调(SFT)以注入缺失知识、稳定训练并促进结构化输出;随后使用RLVR强化结果正确性,精炼决策过程。该方法在多个临床任务中表现优异,涵盖医学计算(MEDCALC)、患者-试验匹配(TREC CLINICAL TRIALS)及疾病诊断(EHRSHOT)。EHRMIND在准确性、可解释性与跨任务泛化方面均取得持续提升,为在医疗场景中应用RLVR提供了实用指导。

原文摘要 · Abstract (English)

We present EHRMIND, a practical recipe for adapting large language models (LLMs) to complex clinical reasoning tasks using reinforcement learning with verifiable rewards (RLVR). While RLVR has succeeded in mathematics and coding, its application to healthcare contexts presents unique challenges due to the specialized knowledge and reasoning required for electronic health record (EHR) interpretation. Our pilot study on the MEDCALC benchmark reveals two key failure modes: (1) misapplied knowledge, where models possess relevant medical knowledge but apply it incorrectly, and (2) missing knowledge, where models lack essential domain knowledge. To address these cases, EHRMIND applies a two-stage solution: a lightweight supervised fine-tuning (SFT) warm-up that injects missing domain knowledge, stabilizes subsequent training, and encourages structured, interpretable outputs; followed by RLVR, which reinforces outcome correctness and refines the model's decision-making. We demonstrate the effectiveness of our method across diverse clinical applications, including medical calculations (MEDCALC), patient-trial matching (TREC CLINICAL TRIALS), and disease diagnosis (EHRSHOT). EHRMIND delivers consistent gains in accuracy, interpretability, and cross-task generalization. These findings offer practical guidance for applying RLVR to enhance LLM capabilities in healthcare settings.

大模型医疗推理强化学习电子病历

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。