将电子病历序列化后训练模型,提升长期医疗关系捕捉能力
Serialized EHR make for good text representations
- 用序列化电子病历数据预训练模型,融合时间与上下文信息
- 在抗生素耐药预测任务中表现优于现有方法,准确率更稳定
- 适合需要理解患者长期诊疗过程的研究者或临床决策系统
医疗领域基础模型的兴起为从大规模临床数据中学习通用表示提供了新路径。然而,现有方法常难以调和电子病历(EHR)的表格化与事件式结构与自然语言模型的序列先验之间的矛盾,导致无法有效捕捉跨就诊记录的纵向依赖关系。本文提出SerialBEHRT,一种领域对齐的基础模型,通过在结构化EHR序列上进一步预训练扩展SciBERT,旨在编码临床事件间的时序与上下文关系,从而生成更丰富的患者表征。我们在抗生素耐药性预测这一具有临床意义的任务上评估其有效性。通过与当前最优的EHR表示策略进行广泛对比,结果表明SerialBEHRT展现出更优且更一致的性能,凸显了在医疗基础模型预训练中引入时序序列化的重要性。
原文摘要 · Abstract (English)
The emergence of foundation models in healthcare has opened new avenues for learning generalizable representations from large scale clinical data. Yet, existing approaches often struggle to reconcile the tabular and event based nature of Electronic Health Records (EHRs) with the sequential priors of natural language models. This structural mismatch limits their ability to capture longitudinal dependencies across patient encounters. We introduce SerialBEHRT, a domain aligned foundation model that extends SciBERT through additional pretraining on structured EHR sequences. SerialBEHRT is designed to encode temporal and contextual relationships among clinical events, thereby producing richer patient representations. We evaluate its effectiveness on the task of antibiotic susceptibility prediction, a clinically meaningful problem in antibiotic stewardship. Through extensive benchmarking against state of the art EHR representation strategies, we demonstrate that SerialBEHRT achieves superior and more consistent performance, highlighting the importance of temporal serialization in foundation model pretraining for healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。