首个英国全国级医疗记录生成模型,用于预测新冠期间及之后的疾病事件。
Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic

- 基于6100万患者长期病历,构建2.43亿参数自回归Transformer模型。
- 零样本预测未来30天住院与死亡风险,支持多种临床编码的精准时间建模。
- 提供可复用的方法框架,揭示大规模医疗数据建模的挑战与路径。
Foresight-England(Foresight-E)是首个全国规模的电子健康记录(EHR)生成基础模型,专为新冠研究设计。该模型在英格兰NHS安全数据环境中从头训练,采用2.43亿参数的Transformer解码器,基于约6100万个体的脱敏纵向病历数据,整合初级/二级护理、死亡登记和新冠数据。训练与验证使用90%数据(5490万人,2018年11月至2022年12月),剩余10%(610万人)用于评估。模型通过自回归方式建模患者时间线,零样本预测任意概念(含约4万种编码),保留ICD-10、OPCS-4和SNOMED CT的临床粒度,并联合表示绝对与相对时间。评估框架涵盖30天新冠住院与死亡风险,按人口统计学特征和疫苗接种状态进行分组分析。为检验对未见未来数据(如2023年)及疫情间接影响的泛化能力,模型在对比逻辑回归与XGBoost基础上测试。由于NHS英格兰已暂停该项目数据访问,当前量化结果暂不可得。本文分享其分词、架构、训练、推理与评估策略,作为构建人群级EHR基础模型的方法论模板与案例研究。
原文摘要 · Abstract (English)
Foresight-England (Foresight-E) is the first national-scale generative foundation model of electronic health records (EHRs), developed as a research pilot strictly for COVID-19 research. We evaluated its ability to model the direct and indirect effects of the pandemic. Trained from scratch entirely within the NHS England Secure Data Environment, Foresight-E is a 243-million-parameter transformer decoder. It was trained and evaluated on de-identified, longitudinal EHRs of approximately 61 million individuals, integrating primary/secondary care, death registrations, and COVID-19 data. Training and validation used a 90% subset (54.9 million) spanning November 2018 to December 2022; the remaining 10% (6.1 million) was held out for evaluation. Foresight-E models patient timelines autoregressively, predicting the next medical event given their prior history. At inference, it operates zero-shot, predicting any concept in its ~40,000-code vocabulary without task-specific training. Our tokenisation scheme retains the clinical granularity of ICD-10, OPCS-4, and SNOMED CT codes, jointly representing absolute and relative timing. We designed an evaluation framework for 30-day COVID-19 hospitalisation and mortality, including subgroup analyses by demographic factors and vaccination status. To assess generalisation to unseen future data and the pandemic's indirect effects, we tested the model on medical events from 2023 (beyond its training period), benchmarking against logistic regression and XGBoost. As detailed in the Project Status section, NHS England has paused access to data for the Foresight-E project, meaning quantitative results are currently unavailable. Instead, we share our strategy for tokenisation, architecture, training, inference, and evaluation as a methodological template and case study in the challenges of building population-scale EHR foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。