用自回归预测下一诊次事件,让电子病历模型无需微调也能精准预测疾病风险。
Foundation Models for Clinical Records at Health System Scale
- 基于患者历史自回归生成下一诊次临床事件,统一处理多类型数据
- 零样本预测痴呆和膝骨关节炎,2年和5年准确率媲美全量微调模型
- 揭示重复事件误判陷阱,提出正则化策略提升评估可靠性
大规模预训练已改变语言及其他数据建模方式,但在结构化电子健康记录(EHR)中的潜力仍待挖掘。本文提出一种针对序列化EHR数据的新型生成式预训练策略,通过预测下一诊次事件实现自回归生成。模型基于患者历史,自动预测多种编码后的临床事件,天然支持异构数据类型的联合预测。此外,我们引入对重复事件的正则化,并指出EHR基础模型评估中的关键缺陷:若不区分新发事件与后续复发事件,重复事件标记会人为抬高性能指标。模型在零样本条件下预测痴呆和膝骨关节炎发病率,时间跨度为2年和5年,表现媲美全量微调的掩码预训练Transformer基线,证明该方法无需任务特定微调即可捕捉复杂临床依赖关系。
原文摘要 · Abstract (English)
Large-scale pretraining has transformed modeling of language and other data types, but its potential remains underexplored in healthcare with structured electronic health records (EHRs). We present a novel generative pretraining strategy for sequential EHR data using next-visit event prediction. Our model learns to autoregressively generate various tokenized clinical events for the next visit based on patient history and inherently handles the joint prediction of heterogeneous data types. Additionally, we introduce regularization on predicting repeated events and highlight a key pitfall in EHR-based foundation model evaluations: repeated event tokens can inflate performance metrics when new onsets are not distinguished from subsequent occurrences. Our model is evaluated via zero-shot prediction for forecasting dementia and knee osteoarthritis incidence within 2 and 5 years, and the model performance rivals a fully fine-tuned masked pretrained Transformer baseline, demonstrating that our approach captures complex clinical dependencies without requiring costly task-specific fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。