构建可预测、可规划的医疗世界模型,提升临床决策可靠性
Beyond Generative AI: World Models for Clinical Prediction, Counterfactuals, and Planning
- 用多模态时序建模学习医疗系统动态,支持多步推演
- 现有系统多达L1-L2级预测与动作条件推理,少数实现反事实推演
- 适合追求可解释、安全的医疗AI研发者与临床决策研究者
医疗AI需具备预测性、可靠性和数据高效性。当前生成模型缺乏物理基础和时序推理能力,难以支撑临床决策。随着语言模型在真实医疗推理中收益递减,世界模型因其能学习反映医疗物理与因果结构的多模态、时序一致、动作依赖表示而受到关注。本文综述了三类医疗世界模型:(i) 医学影像与诊断(如纵向肿瘤模拟、投影-转换建模、JEPA式预测表征学习);(ii) 电子病历中的疾病进展建模(大规模生成事件预测);(iii) 机器人手术与手术规划(动作条件引导与控制)。提出四层能力评估体系:L1时序预测,L2动作条件预测,L3反事实推演支持决策,L4规划/控制。多数系统达L1–L2,L3较少,L4罕见。识别出关键短板:动作空间与安全约束定义不清、干预验证薄弱、多模态状态构建不全、轨迹级不确定性校准不足。本文提出以预测优先的世界模型研究议程,融合生成骨干(Transformer、扩散模型、VAE)与因果/机械基础,实现安全可靠的医疗决策支持。
原文摘要 · Abstract (English)
Healthcare requires AI that is predictive, reliable, and data-efficient. However, recent generative models lack physical foundation and temporal reasoning required for clinical decision support. As scaling language models show diminishing returns for grounded clinical reasoning, world models are gaining traction because they learn multimodal, temporally coherent, and action-conditioned representations that reflect the physical and causal structure of care. This paper reviews World Models for healthcare systems that learn predictive dynamics to enable multistep rollouts, counterfactual evaluation and planning. We survey recent work across three domains: (i) medical imaging and diagnostics (e.g., longitudinal tumor simulation, projection-transition modeling, and Joint Embedding Predictive Architecture i.e., JEPA-style predictive representation learning), (ii) disease progression modeling from electronic health records (generative event forecasting at scale), and (iii) robotic surgery and surgical planning (action-conditioned guidance and control). We also introduce a capability rubric: L1 temporal prediction, L2 action-conditioned prediction, L3 counterfactual rollouts for decision support, and L4 planning/control. Most reviewed systems achieve L1--L2, with fewer instances of L3 and rare L4. We identify cross-cutting gaps that limit clinical reliability; under-specified action spaces and safety constraints, weak interventional validation, incomplete multimodal state construction, and limited trajectory-level uncertainty calibration. This review outlines a research agenda for clinically robust prediction-first world models that integrate generative backbones (transformers, diffusion, VAE) with causal/mechanical foundation for safe decision support in healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。