arXiv:2605.21963cs.LGcs.AI2026-05被引 2

用世界模型预测慢性病患者长期生理变化,比大模型更准。

ChronoMedicalWorld: A Medical World Model for Learning Patient Trajectories from Longitudinal Care Data

  • 构建联合嵌入编码器与动作编码器,融合结构化干预和对话文本。
  • 在2232名肾病患者上,预测每年肌酐清除率误差比GPT-5.5低7.3%。
  • 适合需要长期跟踪干预效果的慢病研究,尤其关注医患沟通影响。

长时程临床模拟——预测患者在特定干预下多年生理变化——是慢病管理的核心,但现有电子健康记录(EHR)模型多为判别式,通用大语言模型在重复干预下易漂移。我们提出 extbf{ChronoMedicalWorld模型(CMWM)},一种基于动作条件的潜在世界模型框架,用于从纵向医疗数据中学习患者轨迹。CMWM结合联合嵌入状态编码器与宽泛动作编码器,支持结构化干预指标和自由文本沟通嵌入,并在六项目标下训练循环潜在转移模块:下一观察监督、下一潜在预测、SIGReg潜在正则化及三项生理感知形状先验(斜率、连续性、大跳跃惩罚)。闭环滚动前缀协议使训练与部署误差一致,模型优化目标即推理时的真实多步误差。以慢性肾病(CKD)为例,基于2,232名肾科患者的数据,该实例在动态50%历史回滚测试中,肌酐清除率(eGFR)预测的平均绝对误差(MAE)为7.384,均方根误差(RMSE)为10.256,优于调优后的GPT-5.5结构化提示基线(MAE 7.964,RMSE 11.069),MAE和RMSE分别降低7.28%和7.35%,提升主要来自患者-健康教练对话部分。该框架不限于肾病,其架构、损失设计与训练协议适用于任何可建模为周期性临床状态与结构化及对话干预交替的慢病。

原文摘要 · Abstract (English)

Long-horizon clinical simulation -- predicting how a patient's physiology evolves over years under specified interventions -- is central to chronic-disease care, yet existing electronic health record (EHR) models are predominantly discriminative, and general-purpose large language models drift under repeated interventions. We propose the \textbf{ChronoMedicalWorld Model (CMWM)}, an action-conditioned latent world-model framework for learning patient trajectories from longitudinal care data. CMWM couples a joint-embedding state encoder with a wide action encoder that admits both structured intervention indicators and free-text communication embeddings, and trains a recurrent latent transition module under a six-term objective: next-observation supervision, next-latent prediction, SIGReg latent regularisation, and three physiology-aware shape priors (slope, continuity, large-jump penalty). A closed-loop rollout-prefix protocol matches training to deployment, so the model is optimised against the same multi-step error it exhibits at inference. As a concrete case study, we instantiate CMWM for annual estimated glomerular filtration rate (eGFR) trajectory forecasting in chronic kidney disease (CKD). On a 2{,}232-patient nephrology cohort, the CKD instantiation achieves a dynamic-50\% history rollout test mean absolute error (MAE) of 7.384 and root-mean-square error (RMSE) of 10.256, against 7.964 and 11.069 for a tuned GPT-5.5 structured-prompting baseline ($-7.28\%$ MAE, $-7.35\%$ RMSE), with the gain dominated by the dialogue portion of patient--health-coach communication. The framework is not CKD-specific: its architecture, loss design, and training protocol apply to any chronic condition that can be cast as periodic clinical state interleaved with structured and conversational interventions.

世界模型慢病预测医疗对话轨迹建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。