arXiv:2509.25591cs.AIcs.CL2025-09被引 6

用未来事件预测增强大模型对病历时间序列的理解能力。

Building the EHR Foundation Model via Next Event Prediction

  • 将病历转化为带时间戳的事件序列,通过自回归训练提升模型时序推理能力。
  • 在肿瘤生存预测和临床诊断任务中,比专用模型高4.6% AUROC,比通用大模型高7.2% C-index。
  • 不仅预测准确,注意力模式还能对应已知疾病发展路径,具备临床可解释性。

电子健康记录(EHR)包含丰富的时序动态,但传统编码方法难以充分捕捉。尽管大语言模型(LLMs)在EHR建模中展现潜力,却在处理序列临床事件和时序依赖方面表现不足。我们提出未来事件预测(NEP)框架,通过在临床事件序列上进行自回归微调,增强LLMs的时序推理能力。将EHR重构为带时间戳的事件链,并预测未来医疗事件,使模型显式建模疾病进展模式与因果关系。在肿瘤生存预测和临床诊断任务中的广泛评估显示,NEP在时序推理任务中优于专用EHR模型4.6% AUROC,优于通用大模型7.2% C-index。分析表明,该方法兼具顶尖预测精度与临床可解释的注意力模式,与已知疾病通路一致。

原文摘要 · Abstract (English)

Electronic Health Records (EHRs) contain rich temporal dynamics that conventional encoding approaches fail to adequately capture. While Large Language Models (LLMs) show promise for EHR modeling, they struggle to reason about sequential clinical events and temporal dependencies. We propose Next Event Prediction (NEP), a framework that enhances LLMs' temporal reasoning through autoregressive fine-tuning on clinical event sequences. By reformulating EHRs as timestamped event chains and predicting future medical events, NEP explicitly models disease progression patterns and causal relationships. Extensive evaluations across oncology survival prediction and clinical diagnosis tasks demonstrate NEP's superiority, outperforming specialized EHR models by 4.6% AUROC and general-purpose LLMs by 7.2% C-index in temporal reasoning tasks. Our analyses reveal dual benefits: state-of-the-art prediction accuracy combined with clinically interpretable attention patterns that align with known disease pathways.

电子病历时序建模大模型临床预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。