arXiv:2608.22622cs.CLcs.AI2026-08

用重症医学专家的推理方式训练大模型,提升跨领域临床判断能力。

Teaching LLMs How ICU Physicians Approach Clinical Reasoning Through OMOP-Aligned Retrieval Improves Reasoning Across Clinical Domains

论文配图:Teaching LLMs How ICU Physicians Approach Clinical Reasoning Through OMOP-Aligned Retrieval Improves Reasoning Across Clinical Domains
图 1 · 摘自论文原文
  • 基于临床医生协作构建重症推理数据集,指导模型进行上下文感知的证据筛选。
  • 在五个临床推理评测中,微调后的模型显著优于基线和通用医疗大模型。
  • 该方法可推广至其他临床任务,适合需复杂推理的医疗AI研发团队。

临床决策依赖于识别相关患者信息以指导诊断与治疗,这在数据密集且动态变化的重症监护室(ICU)尤为困难。大语言模型(LLMs)有望支持此任务,但现有应用和数据集多聚焦表面检索或事实记忆,而非临床医生所实践的归纳与演绎推理。我们假设,通过专家在重症监护中的推理训练,可使大模型具备可泛化的临床推理能力。本文提出ICU-REACT,一个由19名临床医生参与构建的推理数据集,用于训练大模型在ICU中执行信息检索与上下文感知的临床推理。利用该数据集,我们微调了参数规模为8B-70B、涵盖三种模型家族的Clin-REACT模型。在五个临床推理基准测试中,Clin-REACT持续优于其基线模型及开源通用与医疗大模型。性能提升延伸至脚本一致性测试以及下游诊断与治疗任务。结果表明,重症监护领域的专家推理监督可增强更广泛的临床推理能力,但在真实临床使用前仍需前瞻性验证。

原文摘要 · Abstract (English)

Clinical decision-making relies on identifying relevant patient information to guide diagnosis and treatment, a challenge that is especially difficult in the data-dense and rapidly changing intensive care unit (ICU). Large language models (LLMs) could support this task. However, existing applications and datasets mostly emphasize surface-level retrieval or factual recall rather than the inductive and deductive reasoning clinicians practice to select and reason over decision-relevant evidence. We hypothesized that training LLMs on expert ICU reasoning could yield clinical reasoning skills that generalize beyond critical care. Here we introduce ICU-REACT, a reasoning dataset developed with 19 clinicians through a clinician-in-the-loop framework to teach LLMs to perform information retrieval and context-aware clinical reasoning in the ICU. Using ICU-REACT, we fine-tuned Clin-REACT models spanning 8B-70B parameters and three model families. Across five clinical reasoning benchmarks, Clin-REACT consistently outperformed its backbone models and open-source general-purpose and medical LLMs. Gains extended to different tasks including script concordance tests, and downstream diagnosis and treatment tasks. These findings suggest that expert reasoning supervision in critical care can improve broader clinical reasoning, although prospective evaluation is needed before real-world clinical use.

临床推理大模型重症监护知识增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。