arXiv:2411.16346cs.LGstat.ML2024-11中稿 · NeurIPS被引 7

构建首个包含核心治疗变量的大规模重症监护时序数据集,助力跨医院模型迁移。

Towards Foundation Models for Critical Care Time Series

  • 整合多源数据并统一治疗变量,缓解不同医院间的分布偏移问题。
  • 首次构建涵盖关键治疗变量的超大规模重症时序数据集,支持多变量建模。
  • 为重症医疗模型跨院迁移学习提供基准,适合研究分布偏移的学者使用。

通用医疗大语言模型在多个医疗领域取得显著进展,但对重症监护中生命体征、检验结果和治疗记录等住院时序数据的大规模建模仍处于探索阶段。现有数据集规模有限,但通过整合可提升患者多样性并增强模型鲁棒性。为有效利用整合后的数据进行大规模建模,必须解决因治疗政策差异导致的分布偏移问题,关键在于统一各数据集中的治疗变量。本文旨在建立训练大规模多变量时序模型的基础,并为机器学习模型在跨医院迁移学习中应对分布偏移提供基准。我们引入了一个用于序列建模与迁移学习研究的标准化数据集,是首个包含核心治疗变量的超大规模集合。未来计划扩展该数据集,以推动迁移学习发展及可扩展、可泛化的重症医疗模型构建。

原文摘要 · Abstract (English)

Notable progress has been made in generalist medical large language models across various healthcare areas. However, large-scale modeling of in-hospital time series data - such as vital signs, lab results, and treatments in critical care - remains underexplored. Existing datasets are relatively small, but combining them can enhance patient diversity and improve model robustness. To effectively utilize these combined datasets for large-scale modeling, it is essential to address the distribution shifts caused by varying treatment policies, necessitating the harmonization of treatment variables across the different datasets. This work aims to establish a foundation for training large-scale multi-variate time series models on critical care data and to provide a benchmark for machine learning models in transfer learning across hospitals to study and address distribution shift challenges. We introduce a harmonized dataset for sequence modeling and transfer learning research, representing the first large-scale collection to include core treatment variables. Future plans involve expanding this dataset to support further advancements in transfer learning and the development of scalable, generalizable models for critical healthcare applications.

时序建模重症监护迁移学习数据融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。