一个能统一处理病历表示、零样本预测和生成的通用医疗模型
CEHR-XGPT: A Scalable Multi-Task Foundation Model for Electronic Health Records
- 用时间令牌框架建模患者动态病程,支持时序推理
- 在多任务上表现优异,外数据集泛化能力强
- 适合快速开发新医疗应用,无需针对每任务重新训练
电子健康记录(EHR)提供了患者健康的丰富时序视图,对临床决策支持、风险预测和数据驱动的医疗研究具有重要潜力。然而,当前大多数EHR人工智能模型仅针对特定单一任务设计,限制了其通用性和实际应用价值。本文提出CEHR-XGPT,一种面向EHR数据的通用基础模型,统一实现特征表示、零样本预测和合成数据生成三大核心能力。为支持对临床序列的时序推理,该模型引入一种新型基于时间令牌的学习框架,显式将患者动态病程编码至模型结构中。CEHR-XGPT在三个任务上均表现出色,并通过词汇扩展与微调,在外部数据集上展现出良好的泛化能力。其多功能性可实现快速模型开发、队列发现与患者预后预测,无需进行任务特异性再训练。
原文摘要 · Abstract (English)
Electronic Health Records (EHRs) provide a rich, longitudinal view of patient health and hold significant potential for advancing clinical decision support, risk prediction, and data-driven healthcare research. However, most artificial intelligence (AI) models for EHRs are designed for narrow, single-purpose tasks, limiting their generalizability and utility in real-world settings. Here, we present CEHR-XGPT, a general-purpose foundation model for EHR data that unifies three essential capabilities - feature representation, zero-shot prediction, and synthetic data generation - within a single architecture. To support temporal reasoning over clinical sequences, CEHR-XGPT incorporates a novel time-token-based learning framework that explicitly encodes patients' dynamic timelines into the model structure. CEHR-XGPT demonstrates strong performance across all three tasks and generalizes effectively to external datasets through vocabulary expansion and fine-tuning. Its versatility enables rapid model development, cohort discovery, and patient outcome forecasting without the need for task-specific retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。