让大模型学会模拟病人对治疗的反应,提升脓毒症决策安全性
Agentifying Patient Dynamics within LLMs through Interacting with Clinical World Model

- 用临床世界模型模拟病人对液体和升压药的反应
- 在MIMIC-IV数据上比基线更安全且符合指南
- 适合医疗AI研究者和临床决策系统开发者
脓毒症管理需在快速变化的生理状态中做出连续治疗决策。尽管大语言模型(LLMs)具备广泛临床知识并能推理指南,但其本身缺乏对动作条件下的患者动态建模。本文提出SepsisAgent,一种基于世界模型增强的LLM代理,用于脓毒症治疗推荐。SepsisAgent利用学习到的临床世界模型,模拟候选液体-升压药干预下的患者反应,并采用提议-模拟-优化流程后再下达处方。我们首先发现仅使用世界模型无法保证一致的决策性能,因此设计三阶段课程:患者动态监督微调、提议-模拟-优化行为克隆、基于世界模型的代理强化学习。在MIMIC-IV脓毒症轨迹上,SepsisAgent在离策略价值评估中优于所有传统强化学习与基于LLM的基线,同时在指南遵循性与不安全动作指标上表现最佳。进一步分析表明,反复与临床世界模型交互使代理学习到患者演变的规律,即使移除模拟器后仍具实用性。
原文摘要 · Abstract (English)
Sepsis management in the ICU requires sequential treatment decisions under rapidly evolving patient physiology. Although large language models (LLMs) encode broad clinical knowledge and can reason over guidelines, they are not inherently grounded in action-conditioned patient dynamics. We introduce SepsisAgent, a world model-augmented LLM agent for sepsis treatment recommendation. SepsisAgent uses a learned Clinical World Model to simulate patient responses under candidate fluid--vasopressor interventions, and follows a propose--simulate--refine workflow before committing to a prescription. We first show that world-model access alone yields inconsistent LLM decision performance, motivating agent-specific training. We then train SepsisAgent through a three-stage curriculum: patient-dynamics supervised fine-tuning, propose--simulate--refine behavior cloning, and world-model-based agentic reinforcement learning. On MIMIC-IV sepsis trajectories, SepsisAgent outperforms all traditional RL and LLM-based baselines in off-policy value while achieving the best safety profile under guideline adherence and unsafe-action metrics. Further analysis shows that repeated interaction with the Clinical World Model enables the agent to learn regularities in patient evolution, which remain useful even when simulator access is removed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。