多模态心电信号联合建模,提升生理表征学习效果
CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation

- 构建跨模态共享表示,通过延迟对齐匹配不同传感器时间点
- 在25项下游任务中平均提升15.5 AUROC,超越单模态基线
- 适合心电、脉搏、心音等多源医疗信号融合研究者使用
心电图(ECG)、光电容积脉搏波(PPG)和心音图(PCG)提供了同一心动周期的互补视角,但现有心脏基础模型仅针对单一传感模态训练,未充分挖掘跨传感器的共享生理信息。我们提出CardioState-JEPA,一种基于生理感知联合嵌入预测架构的心脏基础模型,可联合学习ECG、PPG和PCG的统一共享表示。该模型将异构波形映射至统一令牌空间,通过单一共享Transformer编码器处理,并通过预测被掩码的潜在心脏状态进行自监督学习,预训练目标聚焦于共享生理而非传感器特定波形外观。为应对电、机械与血流事件间的时序偏移,跨模态预测引入可学习的延迟对齐器,实现对应心动周期时刻的信号匹配。由于同步多传感器数据稀缺,CardioState-JEPA先利用大量单模态数据学习模态内结构,再通过配对数据在潜在心脏时间上对齐各模态。作为冻结编码器在25个涵盖ECG、PPG和PCG的下游任务中评估,其在PPG分类上平均提升8.2 AUROC,在PCG杂音检测上提升18.8 AUROC,ECG分类提升15.5 AUROC,优于最佳自监督信号基线,并在多个ECG基准上达到或超过依赖临床文本或监督标签训练的模型性能。结果表明,异构心脏信号可通过相互监督共同构建单一心脏生理基础模型。
原文摘要 · Abstract (English)
Electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG) provide complementary views of the same cardiac cycle, yet existing cardiac foundation models are trained for a single sensing modality, leaving the shared physiology across sensors unexploited. We introduce CardioState-JEPA, a cardiac foundation model to learn a single shared representation jointly across ECG, PPG, and PCG, built on a physiology-aware joint-embedding predictive architecture. The model maps heterogeneous waveforms into a common token space, processes them with a single shared Transformer encoder, and learns by predicting masked latent cardiac states, placing the pretraining target on shared physiology rather than sensor-specific waveform appearance. To handle the temporal offsets between electrical, mechanical, and hemodynamic events, cross-modal prediction uses a learned delay aligner that matches signals at the corresponding cardiac time. Because synchronized multi-sensor recordings are scarce, CardioState-JEPA first learns within-modality structure from abundant unimodal data and then uses paired data to align modalities in latent cardiac time. Evaluated as a frozen encoder across 25 downstream tasks spanning ECG, PPG, and PCG, our encoder improves average PPG classification by 8.2 AUROC points, PCG murmur detection by 18.8 AUROC points, and ECG classification by 15.5 AUROC points over the best self-supervised signal baseline and matches or exceeds cardiac models trained with privileged clinical text or supervised labels on several ECG benchmarks. These results establish that heterogeneous cardiac signals can mutually supervise a single foundation model of cardiac physiology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。