arXiv:2410.09199cs.LG2024-10被引 4

提出一种高效对比学习方法,用于长时序电子病历数据预训练。

An Efficient Contrastive Unimodal Pretraining Method for EHR Time Series Data

  • 设计负样本估计器,降低对大批量数据的依赖。
  • 在MIMIC-III和eICU上实现更优性能,支持缺失值补全。
  • 适合临床时序数据少样本场景,适用于跨医院迁移应用。

机器学习已革新临床时间序列数据建模。深度神经网络(DNN)可自动学习输入特征与任务间的复杂映射,尤其在重症监护室(ICU)患者长期监测中价值显著。然而,当前最先进的(SOTA)DNN训练方法通常需要大量标注数据,给医疗机构带来高昂成本与时间压力。自监督学习提供替代方案,可在无需昂贵标注的情况下挖掘数据价值。但现有SOTA方法常需大批次数据以达最优性能,增加计算负担,尤其在处理长临床时间序列时面临挑战。为此,本文提出一种针对长临床时间序列数据的高效对比预训练方法。该方法通过负样本估计器实现有效特征提取。我们在标准自监督任务(如线性评估、半监督学习)中验证其有效性,并发现模型具备缺失测量值补全能力,为临床提供更深入的患者状态洞察。实验表明,随着模型规模和测量词表规模扩大,本方法性能持续提升。最后,我们在外部数据集eICU上验证了基于MIMIC-III训练的模型,证明其能学习到可迁移的鲁棒临床信息,适用于不同医疗机构。

原文摘要 · Abstract (English)

Machine learning has revolutionized the modeling of clinical timeseries data. Using machine learning, a Deep Neural Network (DNN) can be automatically trained to learn a complex mapping of its input features for a desired task. This is particularly valuable in Electronic Health Record (EHR) databases, where patients often spend extended periods in intensive care units (ICUs). Machine learning serves as an efficient method for extract meaningful information. However, many state-of-the-art (SOTA) methods for training DNNs demand substantial volumes of labeled data, posing significant challenges for clinics in terms of cost and time. Self-supervised learning offers an alternative by allowing practitioners to extract valuable insights from data without the need for costly labels. Yet, current SOTA methods often necessitate large data batches to achieve optimal performance, increasing computational demands. This presents a challenge when working with long clinical timeseries data. To address this, we propose an efficient method of contrastive pretraining tailored for long clinical timeseries data. Our approach utilizes an estimator for negative pair comparison, enabling effective feature extraction. We assess the efficacy of our pretraining using standard self-supervised tasks such as linear evaluation and semi-supervised learning. Additionally, our model demonstrates the ability to impute missing measurements, providing clinicians with deeper insights into patient conditions. We demonstrate that our pretraining is capable of achieving better performance as both the size of the model and the size of the measurement vocabulary scale. Finally, we externally validate our model, trained on the MIMIC-III dataset, using the eICU dataset. We demonstrate that our model is capable of learning robust clinical information that is transferable to other clinics.

时间序列自监督学习电子病历对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。