用跨模态模型生成可通用的生理信号指纹,实现无重训练的健康监测。
Biosignal Fingerprinting: A Cross-Modal PPG-ECG Foundation Model

- 基于340万对心电与脉搏波数据,训练跨模态自编码器提取通用特征。
- 在7个下游任务中表现优异,最高提升27.7%的分类性能,单模态仍有效。
- 生成隐私保护的个体化心血管状态指纹,适合穿戴设备部署。
心血管疾病仍是全球主要死因,但高诊断价值的心电图(ECG)与普及的可穿戴脉搏波(PPG)之间存在监测鸿沟。为弥合这一差距,本文提出生物信号指纹:一种从跨模态基础模型多模态掩码自编码器(M2AE)中提取的紧凑、可迁移且无需任务重训练的潜在表示。M2AE基于超过340万对配对的ECG与PPG信号训练,采用模态专用编码器与共享瓶颈结构,通过重建和跨模态对比损失联合优化,保留了模态内与跨模态特征。该指纹如同生物特征,以模态无关、隐私保护的形式唯一表征个体心血管状态,可重复用于多种临床任务而无需暴露原始波形或重新训练。在7项下游任务中——包括跨模态重建、五类心血管疾病分类、高血压检测、死亡率预测及人口统计推断——其表现优于或媲美现有领域专用基础模型,如在五分类CVD任务中达到0.974的AUROC,高血压检测达0.877,5项分类任务最高提升27.7%。关键的是,仅需单一模态即可保持强性能,适用于资源受限的单传感器真实可穿戴场景,对临床与消费健康领域的持续心血管监测具有直接意义。
原文摘要 · Abstract (English)
Cardiovascular disease remains the leading cause of global mortality, yet scalable cardiac monitoring is hindered by the gap between diagnostic-rich ECG and ubiquitous wearable PPG. Bridging this gap requires representations that are compact, transferable across modalities and devices, and deployable without task-specific retraining. Here we introduce biosignal fingerprints: compact latent representations of cardiovascular state derived from a cross-modal foundation model, the Multi-modal Masked Autoencoder (M2AE), trained on over 3.4 million paired ECG and PPG signals. M2AE integrates modality-specific encoders with a shared bottleneck and dual decoders, jointly optimized using reconstruction and cross-modal contrastive objectives, yielding generalizable fingerprints that retain intra- and inter-modality features. Like a biometric fingerprint, these representations uniquely encode an individual's cardiovascular state in a modality-agnostic, privacy-preserving form reusable across clinical tasks without exposing raw waveform data or requiring model retraining. Across 7 downstream tasks, spanning cross-modal reconstruction, cardiovascular disease classification, hypertension detection, mortality prediction, and demographic inference, biosignal fingerprints achieve competitive or superior performance compared to leading domain-specialist foundation models in frozen settings, including an AUROC of 0.974 for five-class CVD classification and 0.877 for hypertension detection, with a maximum improvement of 27.7% in AUROC across 5 classification tasks. Critically, strong performance is maintained with only a single modality, enabling deployment in resource-constrained, single-sensor environments typical of real-world wearable monitoring, with direct implications for continuous cardiovascular monitoring across clinical and consumer health settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。