构建可扩展的电子病历基础模型,提升慢病预测能力。
Scaling Electronic Health Record Foundation Models for Population Health Management
- 基于跨机构病历数据预训练,统一编码对齐解决系统异构问题。
- 在11项慢病预测中,美国与台湾队列分别实现70%和40%敏感度,特异性99%。
- 跨系统对齐数据优于单点重复数据,适合资源有限的医疗建模场景。
慢性病管理需要可扩展的方法识别高风险人群,但现有方法依赖碎片化数据且筛查成本高。我们提出电子病历基础模型(EHR-FM),利用来自台湾和美国超过500万患者的数十亿医疗事件进行大规模预训练,采用统一编码对齐框架应对跨系统异构性,并通过IsoFLOP分析揭示其缩放规律,训练出最大达24亿参数的计算最优模型。在11个慢性病预测任务中,该模型表现出强缩放性和泛化能力,优于树模型及通用与生物医学语言模型,在美国和台湾队列中分别实现超过40%和70%的敏感度,特异性均为99%。在EHRShot基准测试中,尽管存在显著分布偏移,仍超越仅在本地数据上训练的前代模型,体现优异的少样本泛化性能。最后,我们证明在数据受限条件下,跨系统对齐数据提供的预训练信号比重复单点数据更有效,凸显对齐对可扩展医疗建模的重要性。分析表明EHR-FM在多种患者分布下稳健,且在ICD编码空间中表现优越。代码将开源。
原文摘要 · Abstract (English)
Population health management requires scalable methods to identify individuals at risk of chronic diseases such as cardiovascular conditions and cancer, yet existing approaches rely on fragmented data and resource-intensive screening. We present Scaling Electronic Health Record Foundation Models for Population Health Management, an Electronic Health Record Foundation Model that performs large-scale chronic disease prediction using cross-site longitudinal medical records. We pretrain Scaling Electronic Health Record Foundation Models for Population Health Management on billions of medical events from over 5 million patients across Taiwan and the United States, leveraging a unified code alignment framework to address cross-system heterogeneity, and characterize its scaling behavior via IsoFLOP analysis, training compute-optimal models up to 2.4B parameters. Across 11 chronic disease prediction tasks, Scaling Electronic Health Record Foundation Models for Population Health Management demonstrates strong scaling and generalization, outperforming tree-based models and both general and biomedical language models, achieving over 40% and 70% sensitivity at 99% specificity in U.S. and Taiwan cohorts, respectively. On the EHRShot benchmark, Scaling Electronic Health Record Foundation Models for Population Health Management surpasses prior EHR foundation models trained on in-site data despite substantial distribution shifts, highlighting strong few-shot generalization. Finally, we show that aligned cross-system data provides more effective pretraining signal than duplicating single-site data under data-limited settings, underscoring the importance of alignment for scalable healthcare modeling. Our analysis demonstrates the robustness of EHR-FM in various patient distributions and the benefits of operating in the ICD code space. The code will be open-sourced.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。