大模型预测患者未来医疗事件,无需微调就能超越专用模型。
Generative Medical Event Models Improve with Scale
- 用1150亿个医疗事件训练生成式模型,自回归预测患者后续健康变化。
- 在78项真实任务中,模型性能随规模提升,无需微调即达顶尖水平。
- 适合临床决策、医疗运营优化,为个性化医疗提供通用框架。
实现大规模个性化医疗需要从纵向患者病程中提取洞见,这些病程可视为一系列医疗事件。基于大规模医疗事件数据预训练的基础模型,是扩展真实世界证据生成并泛化到多样化下游任务的有前景方向。我们使用Epic Cosmos数据集——包含来自310家医疗机构、3亿名唯一患者记录、总计163亿次就诊的匿名长期健康记录——构建了Curiosity系列模型,该系列为仅解码器的Transformer模型,基于1.18亿名患者、1150亿个离散医疗事件(1510亿个标记)进行预训练。我们开展了医疗事件数据最大的缩放规律研究,建立了预训练方法,并揭示了计算量、标记数与模型规模之间的幂律关系。据此,我们训练了一系列计算最优的模型,最大达10亿参数。在给定患者真实历史的基础上,Curiosity可自回归预测下一个医疗事件,以模拟患者健康轨迹。我们评估了78项真实世界任务,包括诊断预测、疾病预后和医疗运营。值得注意的是,对于一个经过通用预训练且采用基于模拟推理的基础模型,Curiosity在这些任务上普遍优于或匹配专用监督模型,且无需任务特定微调或少样本示例。随着模型与预训练规模的增加,其预测能力持续提升。结果表明,Curiosity作为生成式医疗事件基础模型,能有效捕捉复杂临床动态,为支持临床决策、优化医疗运营和改善患者预后提供可扩展、可泛化的框架。
原文摘要 · Abstract (English)
Realizing personalized medicine at scale calls for methods that distill insights from longitudinal patient journeys, which can be viewed as a sequence of medical events. Foundation models pretrained on large-scale medical event data represent a promising direction for scaling real-world evidence generation and generalizing to diverse downstream tasks. Using Epic Cosmos, a dataset with medical events from de-identified longitudinal health records for 16.3 billion encounters over 300 million unique patient records from 310 health systems, we introduce the Curiosity models, a family of decoder-only transformer models pretrained on 118 million patients representing 115 billion discrete medical events (151 billion tokens). We present the largest scaling-law study of medical event data, establishing a methodology for pretraining and revealing power-law scaling relationships for compute, tokens, and model size. Consequently, we pretrained a series of compute-optimal models with up to 1 billion parameters. Conditioned on a patient's real-world history, Curiosity autoregressively predicts the next medical event to simulate patient health timelines. We studied 78 real-world tasks, including diagnosis prediction, disease prognosis, and healthcare operations. Remarkably for a foundation model with generic pretraining and simulation-based inference, Curiosity generally outperformed or matched task-specific supervised models on these tasks, without requiring task-specific fine-tuning or few-shot examples. Curiosity's predictive power consistently improves as the model and pretraining scale. Our results show that Curiosity, a generative medical event foundation model, can effectively capture complex clinical dynamics, providing an extensible and generalizable framework to support clinical decision-making, streamline healthcare operations, and improve patient outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。