用自蒸馏训练多模态生理信号模型,提升危重患者心电/血氧信号质量。
QualityFM: a Multimodal Physiological Signal Foundation Model with Self-Distillation for Signal Quality Challenges in Critically Ill Patients
- 双轨架构+自蒸馏,用高质量信号指导低质量信号建模。
- 在2100万段波形上预训练,跨任务迁移提升诊断准确性。
- 适合临床信号降噪、误报减少与设备研发人员使用。
心电图(ECG)和光电容积脉搏波(PPG)在重症监护室(ICU)和手术室(OR)中广泛应用,但信号质量差、不完整或不一致的问题频发,易导致误报警或诊断错误。现有方法普遍存在泛化能力弱、依赖大量标注数据及跨任务迁移差的问题。为此,我们提出QualityFM,一种新型多模态生理信号基础模型,旨在获得对信号质量的通用理解。模型在包含超过2100万段30秒波形和179,757小时数据的大规模数据集上进行预训练。采用双轨架构处理不同质量的生理信号对,并引入自蒸馏策略:以高质量信号编码器指导低质量信号编码器训练。为高效处理长序列并捕捉局部准周期模式,我们在Transformer模型中集成窗口稀疏注意力机制。此外,复合损失函数结合了编码器输出的直接蒸馏损失与基于功率谱和相位谱的间接重建损失,确保保留信号的频域特性。我们预训练了参数量从960万到3.19亿不等的三个模型,并通过迁移学习在三项临床任务中验证其有效性:室性心动过速误报检测、房颤识别以及从PPG和ECG估计动脉血压(ABP)。
原文摘要 · Abstract (English)
Photoplethysmogram (PPG) and electrocardiogram (ECG) are commonly recorded in intesive care unit (ICU) and operating room (OR). However, the high incidence of poor, incomplete, and inconsistent signal quality, can lead to false alarms or diagnostic inaccuracies. The methods explored so far suffer from limited generalizability, reliance on extensive labeled data, and poor cross-task transferability. To overcome these challenges, we introduce QualityFM, a novel multimodal foundation model for these physiological signals, designed to acquire a general-purpose understanding of signal quality. Our model is pre-trained on an large-scale dataset comprising over 21 million 30-second waveforms and 179,757 hours of data. Our approach involves a dual-track architecture that processes paired physiological signals of differing quality, leveraging a self-distillation strategy where an encoder for high-quality signals is used to guide the training of an encoder for low-quality signals. To efficiently handle long sequential signals and capture essential local quasi-periodic patterns, we integrate a windowed sparse attention mechanism within our Transformer-based model. Furthermore, a composite loss function, which combines direct distillation loss on encoder outputs with indirect reconstruction loss based on power and phase spectra, ensures the preservation of frequency-domain characteristics of the signals. We pre-train three models with varying parameter counts (9.6 M to 319 M) and demonstrate their efficacy and practical value through transfer learning on three distinct clinical tasks: false alarm of ventricular tachycardia detection, the identification of atrial fibrillation and the estimation of arterial blood pressure (ABP) from PPG and ECG signals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。