arXiv:2605.09173cs.LGcs.AI2026-05

用分层自监督学习,从数百万小时可穿戴设备波形中挖掘长期生理规律。

WavesFM: Hierarchical Representation Learning for Longitudinal Wearable Sensor Waveforms

论文配图:WavesFM: Hierarchical Representation Learning for Longitudinal Wearable Sensor Waveforms
图 1 · 摘自论文原文
  • 分两阶段训练:先学短段波形特征,再建模多日序列变化
  • 在680万小时数据上预训练,跨58项任务表现更优
  • 适合做长期健康监测与疾病早期预警的研究者

可穿戴传感器可在自由生活状态下持续采集高分辨率生理波形(如光电容积脉搏波、加速度计信号)。但因采样频率高、多模态依赖强、序列极长(如数周记录)且标注数据稀缺,从中推断健康表型面临巨大挑战。现有自监督学习方法通常要么仅关注短段波形的形态特征而忽略长期动态,要么基于人工设计的粗粒度特征(如心率、步数)建模行为模式,放弃原始波形中的细微预测信号。为此,我们提出WavesFM,一种用于纵向生理数据的双阶段自监督学习基础模型。首先在短波形片段上预训练段级编码器提取局部嵌入;随后在多日时间尺度上训练时序编码器建模这些嵌入序列。该分层策略克服了高分辨率长序列数据的计算复杂性,使模型同时捕捉局部信号语义与昼夜及跨日变化规律。第一阶段在超过680万小时(N=32.4万人)的记录上预训练,第二阶段在530万小时(N=1万人)上训练。WavesFM在涵盖人口统计、生活方式、健康状况和用药等58项任务中表现优异。

原文摘要 · Abstract (English)

Wearable sensors enable the continuous acquisition of high-resolution physiological waveforms, such as photoplethysmography and accelerometry, under free-living conditions. However, inferring health-related phenotypes from these signals presents significant challenges due to high sampling frequencies, multimodal dependencies, and extreme sequence lengths (e.g., weeks of recordings), compounded by a scarcity of ground-truth labels. To address these challenges, existing self-supervised learning (SSL) methodologies typically follow two paradigms: (1) learning rich morphological representations from short waveform segments while collapsing longitudinal dynamics through simple aggregation, or (2) modeling behavioral patterns from coarse, hand-crafted features (e.g. heart rate, step counts) spanning longer horizons but foregoing subtle, predictive signatures in raw waveforms. To bridge this gap, we propose WavesFM, a foundation model utilizing a two-stage SSL framework for longitudinal physiological data. Specifically, we decompose the learning problem into two stages: first, a segment-level encoder is pretrained to extract local embeddings from short waveforms; subsequently, a temporal encoder is trained to model the sequence of these embeddings across a multi-day horizon. This hierarchical approach overcomes the computational complexity of high-resolution, long-sequence data, allowing the overall model to capture both local signal semantics and the complex circadian and inter-day variations governing physiological dynamics. Pretrained on over 6.8M hours (N=324k individuals) of recordings for the first stage and 5.3M hours (N=10k) for the second stage, WavesFM demonstrates superior performance across 58 diverse tasks spanning demographics, lifestyle, health conditions, and medications.

可穿戴设备自监督学习生理信号长序列建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。