用大模型模拟人体多器官生理,可生成真实病程数据用于临床研究。
Organ-Agents: Virtual Human Physiology Simulator via LLMs

- 每个器官由独立大模型代理驱动,通过时序数据训练与强化协调机制。
- 在4509例患者上实现各系统误差小于0.16,关键事件时间同步准确。
- 支持治疗方案对比仿真,适合重症医学研究与精准医疗探索。
大语言模型(LLMs)的进展为复杂生理系统模拟带来新可能。我们提出Organ-Agents,一种基于大模型代理的多智能体框架,用于模拟人体生理。每个模拟器对应特定系统(如心血管、肾脏、免疫)。训练包括基于系统特异性时序数据的监督微调,以及利用动态参考选择和错误纠正的强化引导协调。我们整合了7,134名脓毒症患者与7,895名对照者的数据,生成覆盖9个系统、125个变量的高分辨率轨迹。在4,509例保留患者上,各系统均方误差(MSE)低于0.16,且在不同SOFA严重程度分层中表现稳健。在两家医院共22,689例ICU患者上外部验证显示,分布偏移下性能略有下降但模拟保持稳定。该系统能准确重现关键多系统事件(如低血压、高乳酸血症、低氧血症),具有合理的时间顺序与阶段进展。15位重症医生评估确认其现实性与生理合理性(平均李克特评分3.9和3.7)。此外,该系统可进行替代治疗策略的反事实模拟,生成与真实患者匹配的轨迹及APACHE II评分。在下游早期预警任务中,基于合成数据训练的分类器AUROC下降不足0.04,表明决策相关模式得以保留。这些结果表明,Organ-Agents是可信、可解释、通用的数字孪生系统,适用于精准诊断、治疗模拟与假说检验。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have enabled new possibilities in simulating complex physiological systems. We introduce Organ-Agents, a multi-agent framework that simulates human physiology via LLM-driven agents. Each Simulator models a specific system (e.g., cardiovascular, renal, immune). Training consists of supervised fine-tuning on system-specific time-series data, followed by reinforcement-guided coordination using dynamic reference selection and error correction. We curated data from 7,134 sepsis patients and 7,895 controls, generating high-resolution trajectories across 9 systems and 125 variables. Organ-Agents achieved high simulation accuracy on 4,509 held-out patients, with per-system MSEs <0.16 and robustness across SOFA-based severity strata. External validation on 22,689 ICU patients from two hospitals showed moderate degradation under distribution shifts with stable simulation. Organ-Agents faithfully reproduces critical multi-system events (e.g., hypotension, hyperlactatemia, hypoxemia) with coherent timing and phase progression. Evaluation by 15 critical care physicians confirmed realism and physiological plausibility (mean Likert ratings 3.9 and 3.7). Organ-Agents also enables counterfactual simulations under alternative sepsis treatment strategies, generating trajectories and APACHE II scores aligned with matched real-world patients. In downstream early warning tasks, classifiers trained on synthetic data showed minimal AUROC drops (<0.04), indicating preserved decision-relevant patterns. These results position Organ-Agents as a credible, interpretable, and generalizable digital twin for precision diagnosis, treatment simulation, and hypothesis testing in critical care.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。