arXiv:2505.02722cs.AIcs.LG2025-05中稿 · MLHC 2026被引 2

用全国脓毒症数据微调大模型,提升临床推理能力

Enhancing LLMs' Clinical Reasoning with Real-World Data from a Nationwide Sepsis Registry

  • 从脓毒症注册库构建推理题,用强化学习微调Phi-4
  • 在同领域测试中表现优异,专家评估认可推理质量
  • 能力可推广至不同任务、患者群体及其它疾病场景

尽管大型语言模型在通用领域展现出强大的推理能力,但在真实临床实践中应用仍受限,这可能源于训练时缺乏真实临床数据——由于隐私问题,这类数据通常未被纳入。为此,我们提出通过利用真实世界临床数据来增强大模型的临床推理能力。我们基于全国脓毒症注册库构建了高推理强度的问题,并使用强化学习对Phi-4进行微调,得到C-Reason模型。C-Reason在同域测试集上表现出色,不仅在量化指标上优于基线,且通过专家评估验证了其推理质量。此外,其增强的推理能力还泛化到不同任务、患者群体的脓毒症数据集,以及抗生素使用开放式咨询任务,甚至其他疾病场景。未来研究应致力于使用大规模、多病种临床数据训练大模型,以发展更强大、通用的临床推理系统。

原文摘要 · Abstract (English)

Although large language models (LLMs) have demonstrated impressive reasoning capabilities across general domains, their effectiveness in real-world clinical practice remains limited. This is likely due to their insufficient exposure to real-world clinical data during training, as such data is typically not included due to privacy concerns. To address this, we propose enhancing the clinical reasoning capabilities of LLMs by leveraging real-world clinical data. We constructed reasoning-intensive questions from a nationwide sepsis registry and fine-tuned Phi-4 on these questions using reinforcement learning, resulting in C-Reason. C-Reason exhibited strong clinical reasoning capabilities on the in-domain test set, as evidenced by both quantitative metrics and expert evaluations. Furthermore, its enhanced reasoning capabilities generalized to a sepsis dataset involving different tasks and patient cohorts, an open-ended consultations on antibiotics use task, and other diseases. Future research should focus on training LLMs with large-scale, multi-disease clinical datasets to develop more powerful, general-purpose clinical reasoning models.

临床推理大模型脓毒症数据微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。