arXiv:2604.22780eess.SPcs.AI2026-04被引 11

用自监督学习从少量标注数据中精准估计生理参数,提升非侵入式监测效果。

A General Framework for Generative Self-supervised Learning in Non-invasive Estimation of Physiological Parameters Using Photoplethysmography

  • 设计跨时序融合生成锚点任务,捕捉心率信号的全局与局部特征。
  • 在仅10%训练数据下,均方根误差降低2.49%,超越现有方法。
  • 适合医疗健康领域研究人员,尤其关注无创生理监测的场景。

将生理参数标签与大规模光电容积脉搏波(PPG)数据对齐进行深度学习面临挑战且成本高昂。尽管自监督表示学习(SSRL)可应对标注数据有限的问题,但仍需从大量未标注数据中学习鲁棒的共享表示,并整合上下文线索以获得差异性表征。为此,提出一种生成式自监督学习框架TS2TC,利用时间域、频谱图域及混合域探索并融合PPG的独特特征,实现通用且非侵入式的生理参数估计。设计了一种名为跨时序融合生成锚点(CTFGA)的预训练任务,建模时间依赖性并在粗粒度上重构独立片段,从而实现鲁棒的全局特征提取与局部上下文表示。框架还引入了具有不同频率尺度和导数阶次的子信号,反映血流动力学特性,促进多语义层次的共享表征学习。其次,提出受认知启发的双过程迁移(DPT)策略,包含先验依赖的自主过程与后验观测推理过程,兼顾共享与特定表示的优势。TS2TC在混合域中引入双线性时-频谱融合方法,对齐不同域的潜在表示,建立多源信息间的细粒度上下文交互。在多个生理参数估计任务上的大量实验表明,CTFGA与DPT联合性能显著优于标准生成式学习。在仅使用10%训练数据的情况下,平均均方根误差(RMSE)相比最先进方法降低2.49%。

原文摘要 · Abstract (English)

Aligning physiological parameter labels with large-scale photoplethysmographic (PPG) data for deep learning is challenging and resource-intensive. While self-supervised representation learning (SSRL) can handle limited annotated data, the challenge lies in learning robust shared representations from vast unlabeled data and integrating contextual cues to learn distinctive representations. To alleviate these challenges, a generative SSRL framework TS2TC is proposed to utilize the temporal, spectrogram, and temporal-spectrogram mixed domains to explore and incorporate the unique features of PPG for universal and noninvasive physiological parameter estimation. A pretext task named Cross-Temporal Fusion Generative Anchor (CTFGA) is designed, modeling temporal dependencies and reconstructing independent segments at a coarse level to provide robust global feature extraction and local contextual representation. The framework includes sub-signals from PPG with diverse frequency scales and order derivatives reflecting hemodynamics to facilitate learning shared representations at varying semantic levels. Secondly, a cognitive-inspired dual-process transfer (DPT) strategy is formulated, consisting of prior-dependent autonomous processes and posterior observation reasoning processes, to leverage the independent and integrated advantages of shared and specific representations. TS2TC introduces a bilinear temporal-spectrogram fusion method in the mixed domain, aligning latent representations from different domains and establishing fine-grained contextual interactions across multiple sources of information. Extensive experiments on physiological parameter estimation tasks showed that the joint performance of CTFGA and DPT outperforms standard generative learning significantly. TS2TC achieved an average 2.49\% improvement in RMSE over state-of-the-art estimation methods with only 10\% training data.

自监督学习生理参数估计光电容积脉搏波非侵入式监测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。