基于170万患者数据训练的心脏健康多模态大模型,可跨设备场景精准分析心电与光电信号。
Sensing Cardiac Health Across Scenarios and Devices: A Multi-Modal Foundation Model Pretrained on Heterogeneous Data from 1.7 Million Individuals
- 用Transformer和掩码预训练学习跨模态统一表征
- 在12导联、单导联及仅心电/光电信号下均表现优异
- 适合临床诊断、生命体征监测等多场景应用
心脏生物信号(如心电图ECG和光电容积脉搏波PPG)对心血管疾病的诊断、预防与管理至关重要,广泛应用于各类临床任务。传统深度学习方法通常依赖同质数据集和专用模型,限制了其在多样临床环境和采集协议下的鲁棒性与泛化能力。本研究提出心脏感知基础模型(CSFM),利用先进的Transformer架构与生成式掩码预训练策略,从大规模异构健康记录中学习统一表征。模型在包含约170万个体的多个大型数据集(MIMIC-III-WDB、MIMIC-IV-ECG、CODE)的多模态数据上预训练,涵盖心脏信号及其对应的临床或机器生成文本报告。实验表明,CSFM提取的嵌入向量在多种心脏感知场景中均能有效作为特征提取器,并支持跨输入配置与传感器模态的无缝迁移学习。在诊断任务、人口统计信息识别、生命体征测量、临床结局预测及心电图问答等多项评估中,CSFM持续优于传统单模态单任务方法。尤其在标准12导联系统到单导联设置,以及仅有ECG、仅有PPG或两者组合的场景下表现稳健。这些发现凸显了CSFM作为全面、可扩展的心脏监测解决方案的潜力。
原文摘要 · Abstract (English)
Cardiac biosignals, such as electrocardiograms (ECG) and photoplethysmograms (PPG), are of paramount importance for the diagnosis, prevention, and management of cardiovascular diseases, and have been extensively used in a variety of clinical tasks. Conventional deep learning approaches for analyzing these signals typically rely on homogeneous datasets and static bespoke models, limiting their robustness and generalizability across diverse clinical settings and acquisition protocols. In this study, we present a cardiac sensing foundation model (CSFM) that leverages advanced transformer architectures and a generative, masked pretraining strategy to learn unified representations from vast, heterogeneous health records. Our model is pretrained on an innovative multi-modal integration of data from multiple large-scale datasets (including MIMIC-III-WDB, MIMIC-IV-ECG, and CODE), comprising cardiac signals and the corresponding clinical or machine-generated text reports from approximately 1.7 million individuals. We demonstrate that the embeddings derived from our CSFM not only serve as effective feature extractors across diverse cardiac sensing scenarios, but also enable seamless transfer learning across varying input configurations and sensor modalities. Extensive evaluations across diagnostic tasks, demographic information recognition, vital sign measurement, clinical outcome prediction, and ECG question answering reveal that CSFM consistently outperforms traditional one-modal-one-task approaches. Notably, CSFM exhibits robust performance across multiple ECG lead configurations from standard 12-lead systems to single-lead setups, and in scenarios where only ECG, only PPG, or a combination thereof is available. These findings highlight the potential of CSFM as a versatile and scalable solution, for comprehensive cardiac monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。