arXiv:2601.21031cs.LGcs.AI2026-01被引 3

用统计先验指导生成掩码,提升心率信号模型的抗噪与泛化能力。

SIGMA-PPG: Statistical-prior Informed Generative Masking Architecture for PPG Foundation Model

  • 引入统计先验驱动的对抗性掩码机制,防止模型过拟合噪声。
  • 通过向量量化约束语义一致性,使相似波形映射到相同编码索引。
  • 在12万小时数据上预训练,12项下游任务均优于现有方法。

当前光电容积脉搏波(PPG)基础模型面临信号固有的冗余与噪声挑战。标准掩码建模常导致平凡解,而对比方法缺乏形态精确性。为此,我们提出统计先验引导的生成掩码架构(SIGMA-PPG),其包含由强化学习驱动的教师网络,利用统计先验生成具有挑战性的学习路径,避免对噪声的过拟合。同时,通过向量量化引入语义一致性约束,确保生理相似的波形(即使受采集伪影或微小扰动影响)映射至共享编码索引,增强代码本语义密度并消除冗余特征结构。在超过12万小时数据上预训练后,SIGMA-PPG在12个多样化下游任务中的平均表现优于五种先进基线方法。代码已公开于 https://github.com/ZonghengGuo/SigmaPPG。

原文摘要 · Abstract (English)

Current foundation model for photoplethysmography (PPG) signals is challenged by the intrinsic redundancy and noise of the signal. Standard masked modeling often yields trivial solutions while contrastive methods lack morphological precision. To address these limitations, we propose a Statistical-prior Informed Generative Masking Architecture (SIGMA-PPG), a generative foundation model featuring a Prior-Guided Adversarial Masking mechanism, where a reinforcement learning-driven teacher leverages statistical priors to create challenging learning paths that prevent overfitting to noise. We also incorporate a semantic consistency constraint via vector quantization to ensure that physiologically identical waveforms (even those altered by recording artifacts or minor perturbations) map to shared indices. This enhances codebook semantic density and eliminates redundant feature structures. Pre-trained on over 120,000 hours of data, SIGMA-PPG achieves superior average performance compared to five state-of-the-art baselines across 12 diverse downstream tasks. The code is available at https://github.com/ZonghengGuo/SigmaPPG.

PPG生成模型先验引导向量量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。