arXiv:2607.15447cs.LG2026-07

用大模型对齐病历事件与时间序列,提升重症监护预测性能。

LLM4EHR: Aligning Clinical Time Series with Medical Event Sequences via Large Language Models

论文配图:LLM4EHR: Aligning Clinical Time Series with Medical Event Sequences via Large Language Models
图 1 · 摘自论文原文
  • 通过对比学习对齐病历事件与时间序列数据
  • 在多个下游任务中表现优于现有方法
  • 支持少样本迁移,适合新患者群体应用

近年来,临床机器学习在重症监护室(ICU)的预后预测中,从专用监督模型转向基础模型,利用现代表示学习方法。基础模型在混合复杂临床数据模态上预训练,适用于多种下游任务。现有工作多使用电子健康记录(EHR)提供丰富的患者观测数据来训练临床基础模型,但未充分探索病历事件与时间序列(TS)观测之间的共享时间结构。这一局限可能导致模型鲁棒性与适应性不足,影响下游任务表现。为充分挖掘这种时间结构,我们提出 LLM4EHR,一种在 ICU EHR 数据上训练的新临床基础模型。结合领域适配的大语言模型与变换器时间序列编码器,通过时间对齐病历事件与时间序列进行预训练。为此,我们提出一种正则化对比目标,学习以病历事件嵌入为条件的时间序列表示。消融实验证明,LLM4EHR 学得的病历时间序列嵌入显著提升多种下游临床任务性能。此外,我们实证表明,该模型能通过少样本(k-shot)适配将可迁移的临床时间序列嵌入部署至新队列。这些发现推动了更通用、高性能临床基础模型的发展。

原文摘要 · Abstract (English)

Recent research in clinical machine learning, focusing on outcome predictions in intensive care unit (ICU), has shifted from bespoke supervised models to foundation models, utilising modern representation learning methods. Here, foundation models are pre-trained on mixtures of complex clinical data modalities, useful for various downstream tasks. Existing works often utilise Electronic Health Records (EHR) to provide rich and diverse patient observations to train clinical foundation models. However, existing methods do not sufficiently explore the shared temporal structures between clinical events and time series (TS) observations recorded in EHRs. This limitation potentially leads to less robust and adaptive clinical foundation models, resulting in reduced performance on downstream tasks. To fully exploit this temporal structure, we propose LLM4EHR, a new clinical foundation model trained on ICU EHR data. Combining domain adapted large language models with a transformer TS encoder, we pre-trained LLM4EHR by temporally aligning the EHR events and TS. For this, we propose a regularised contrastive objective to learn robust EHR TS representations conditioned on EHR event embeddings produced by the domain adapted LLM. Supported by an ablation study, we find that learnt EHR TS embeddings from LLM4EHR improve performance on various downstream clinical tasks with competitive performance. Further, we empirically demonstrate that LLM4EHR learns transferable clinical TS embeddings that can be deployed to new cohorts via k-shot adaptation. These findings provide a step towards building more generalisable and performant clinical foundation models.

临床模型时间序列大模型EHR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。