用秒级标签提升可穿戴设备睡眠分期准确率
Rethinking PPG-based Sleep Staging: Datasets, Metrics, and Benchmarks

- 基于隐半马尔可夫模型将30秒标签扩展为秒级标注
- 在MESA数据集上使分期准确率提升3.7%~5.7个百分点
- 方法对不同人群和标注协议均具迁移性,适合可穿戴睡眠研究
自动化睡眠分期将离散阶段标签分配给整晚记录中的连续时间片段,传统上每个时间窗至少30秒,反映临床评分标准的最小时间分辨率。可穿戴光电容积脉搏波(PPG)作为实验室多导睡眠图(PSG)的便携替代方案受到持续关注,后者依赖脑电图(EEG)等在非临床环境不实用的模态。然而,基于PPG的分期仍显著落后于基于EEG的方法,我们认为这主要源于信号与任务之间的不匹配。在稳定阶段内,PPG的阶段间特征差异比EEG更细微;但在阶段边界处,PPG的主要心血管特征——心率变异性与脉搏形态——在数秒内发生剧烈变化。传统的每30秒一个标签的做法因此抑制了集中在边界附近的特征。我们通过两步解决这一差距:首先,基于隐半马尔可夫模型开发标签扩展流程,将粗粒度的时间窗标签转换为秒级标注;为验证这些扩展标签是否足以作为下游监督信号,我们在独立专家评审数据集上进行验证,并通过一个独立的睡眠-清醒任务(其标签不受扩展流程影响)进一步检验。其次,利用所得秒级监督在MESA数据集上,使四种架构各异的基线模型在四类阶段划分任务上的准确率相对于原始30秒标签提升了3.7–5.7个百分点,零样本迁移评估在CFS数据集上显示该增益在队列和标注协议转移下依然成立。
原文摘要 · Abstract (English)
Automated sleep staging assigns discrete stage labels to successive time epochs throughout an overnight recording; conventionally each window spans at least 30 seconds, reflecting the minimum temporal resolution of the clinical scoring standard. Wearable photoplethysmography (PPG) has attracted sustained interest as an ambulatory alternative to laboratory-based polysomnography, which relies on electroencephalography (EEG) and other recording modalities that are impractical outside clinical environments. Yet PPG-based staging trails EEG-based methods by a substantial margin, and we argue this gap largely reflects a mismatch between signal and task. Within a stable stage, PPG's inter-stage feature differences are more subtle than those in EEG; yet at stage boundaries, PPG's principal cardiovascular features, heart rate variability and pulse morphology, shift sharply within seconds. The conventional practice of assigning one label to each 30-second epoch therefore suppresses feature that is concentrated near boundaries. We address this gap in two steps. First, we develop a label expansion pipeline based on Hidden Semi-Markov Models that converts coarse epoch labels into sec-level annotations. To assess whether these expanded labels are reliable enough for downstream supervision, we validate them on a separate expert-reviewed dataset and through an auxiliary sleep-wake task whose labels are independent of the expansion pipeline. Second, we use the resulting sec-level supervision on MESA to improve conventional four-class epoch-level staging across four architecturally diverse baselines by 3.7--5.7\,pp in accuracy against the original epoch labels, with supplementary zero-shot evaluation on CFS showing that the transfer benefit persists under cohort and annotation-protocol shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。