arXiv:2502.19593cs.LGcs.AI2025-02中稿 · AAAI被引 1

用预训练模型ICU-BERT提升重症医疗数据表征能力

Improving Representation Learning of Complex Critical Care Data with ICU-BERT

  • 基于MIMIC-IV数据用多任务训练Transformer,学习复杂临床数据表示
  • 在5项任务和4个新数据集上表现优于或持平现有基准
  • 融合结构化与非结构化数据,适合临床决策支持场景

真实世界临床数据(如重症监护室数据)具有多变量、异步特点,传统AI系统常假设数据规则且特征独立,依赖有限数据范围和人工特征工程,未充分挖掘生成式AI潜力。本文提出ICU-BERT,一种基于Transformer的预训练模型,利用MIMIC-IV数据库通过多任务方案学习复杂重症数据的鲁棒表征,仅需最少预处理。ICU-BERT采用多标记输入策略,结合生物医学大语言模型的密集嵌入,学习通用的复杂多变量重症数据表示。初步评估涵盖5项任务及4个额外重症数据集,结果表明:通过微调,ICU-BERT性能可媲美或超越当前主流基准。通过整合结构化与非结构化数据,该模型推动了基础模型在医学信息学中的应用,为多种临床决策支持场景提供可适配解决方案。

原文摘要 · Abstract (English)

The multivariate, asynchronous nature of real-world clinical data, such as that generated in Intensive Care Units (ICUs), challenges traditional AI-based decision-support systems. These often assume data regularity and feature independence and frequently rely on limited data scopes and manual feature engineering. The potential of generative AI technologies has not yet been fully exploited to analyze clinical data. We introduce ICU-BERT, a transformer-based model pre-trained on the MIMIC-IV database using a multi-task scheme to learn robust representations of complex ICU data with minimal preprocessing. ICU-BERT employs a multi-token input strategy, incorporating dense embeddings from a biomedical Large Language Model to learn a generalizable representation of complex and multivariate ICU data. With an initial evaluation of five tasks and four additional ICU datasets, ICU-BERT results indicate that ICU-BERT either compares to or surpasses current performance benchmarks by leveraging fine-tuning. By integrating structured and unstructured data, ICU-BERT advances the use of foundational models in medical informatics, offering an adaptable solution for clinical decision support across diverse applications.

重症监护表征学习预训练模型医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。