让医疗模型学会理解病程演变关系,提升预测与诊断能力。
Temporal Entailment Pretraining for Clinical Language Models over EHR Data
- 用时间顺序的医案片段训练模型判断前后状态的逻辑关系
- 在MIMIC-IV数据上实现临床问答、预警预测等任务最优性能
- 适合关注病程建模与时序推理的临床AI研究者
临床语言模型通过在出院小结和病历等专有语料上预训练,已在下游任务中取得优异表现。然而,多数方法将电子健康记录视为静态文档,忽视了患者病程随时间演进且因果交织的本质。本文提出一种新型的时序蕴含预训练目标,将EHR片段构造成时间有序的句子对,训练模型判断后一状态是否被前一状态蕴含、矛盾或中立。通过这一结构化时序预训练任务,模型学会隐式进行临床推理,显著提升在预测与诊断任务中的泛化能力。我们在MIMIC-IV大规模语料上进行预训练,在时序临床问答、早期预警预测和疾病进展建模任务上均达到当前最优效果。
原文摘要 · Abstract (English)
Clinical language models have achieved strong performance on downstream tasks by pretraining on domain specific corpora such as discharge summaries and medical notes. However, most approaches treat the electronic health record as a static document, neglecting the temporally-evolving and causally entwined nature of patient trajectories. In this paper, we introduce a novel temporal entailment pretraining objective for language models in the clinical domain. Our method formulates EHR segments as temporally ordered sentence pairs and trains the model to determine whether a later state is entailed by, contradictory to, or neutral with respect to an earlier state. Through this temporally structured pretraining task, models learn to perform latent clinical reasoning over time, improving their ability to generalize across forecasting and diagnosis tasks. We pretrain on a large corpus derived from MIMIC IV and demonstrate state of the art results on temporal clinical QA, early warning prediction, and disease progression modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。