用深度学习自动判读帕金森病患者睡眠阶段,准确率超人工。
Fully-automated sleep staging: multicenter validation of a generalizable deep neural network for Parkinson's disease and isolated REM sleep behavior disorder
- 基于多中心数据预训练+微调,模型适应神经退行性疾病异常脑电。
- 在独立测试集上,睡眠分期一致性达κ=0.69,显著优于原始模型。
- 引入置信度阈值后,快速识别正确REM睡眠,95%患者保留足够时长。
孤立性快速眼动期睡眠行为障碍(iRBD)是帕金森病(PD)的重要前驱标志,视频多导睡眠图(vPSG)仍是诊断金标准。但神经退行性疾病中因脑电异常和睡眠碎片化,人工分期困难,成为大规模筛查技术部署的瓶颈。本文将U-Sleep深度神经网络适配于PD与iRBD患者的通用睡眠分期任务。基于包含19,236例多中心非神经退行性疾病PSG数据的PUB数据集预训练模型,再在吕德贝克基金会帕金森病研究中心(PACE)和科隆-波恩队列(CBC)两个中心的研究数据集(112例PD、138例iRBD、89例年龄匹配对照)上进行微调。最终模型在丹麦睡眠医学中心(DCSM)独立数据集(81例PD、36例iRBD、87例睡眠门诊对照)上评估。部分与人类评分者一致性低(Cohen's κ < 0.6)的PSG由第二名盲评人重新打分,以识别分歧原因。通过引入置信度阈值优化快速眼动期(REM)睡眠判定。预训练模型在PUB上平均κ=0.81,直接应用于PACE/CBC时降至κ=0.66。微调后模型在PACE/CBC上κ=0.74(p<0.001 vs. 预训练模型)。在DCSM中,平均κ从0.60升至0.64(p<0.001),中位κ从0.64升至0.69(p<0.001)。人与人之间的评分一致性研究显示,模型与初始评分者不一致的记录,其人工评分间一致性也较低。应用置信度阈值后,正确识别的REM睡眠段比例从85%提升至95.5%,同时95%受试者仍保留超过5分钟的有效REM睡眠时间。
原文摘要 · Abstract (English)
Isolated REM sleep behavior disorder (iRBD) is a key prodromal marker of Parkinson's disease (PD), and video-polysomnography (vPSG) remains the diagnostic gold standard. However, manual sleep staging is particularly challenging in neurodegenerative diseases due to EEG abnormalities and fragmented sleep, making PSG assessments a bottleneck for deploying new RBD screening technologies at scale. We adapted U-Sleep, a deep neural network, for generalizable sleep staging in PD and iRBD. A pretrained U-Sleep model, based on a large, multisite non-neurodegenerative dataset (PUB; 19,236 PSGs across 12 sites), was fine-tuned on research datasets from two centers (Lundbeck Foundation Parkinson's Disease Research Center (PACE) and the Cologne-Bonn Cohort (CBC); 112 PD, 138 iRBD, 89 age-matched controls. The resulting model was evaluated on an independent dataset from the Danish Center for Sleep Medicine (DCSM; 81 PD, 36 iRBD, 87 sleep-clinic controls). A subset of PSGs with low agreement between the human rater and the model (Cohen's $κ$ < 0.6) was re-scored by a second blinded human rater to identify sources of disagreement. Finally, we applied confidence-based thresholds to optimize REM sleep staging. The pretrained model achieved mean $κ$ = 0.81 in PUB, but $κ$ = 0.66 when applied directly to PACE/CBC. By fine-tuning the model, we developed a generalized model with $κ$ = 0.74 on PACE/CBC (p < 0.001 vs. the pretrained model). In DCSM, mean and median $κ$ increased from 0.60 to 0.64 (p < 0.001) and 0.64 to 0.69 (p < 0.001), respectively. In the interrater study, PSGs with low agreement between the model and the initial scorer showed similarly low agreement between human scorers. Applying a confidence threshold increased the proportion of correctly identified REM sleep epochs from 85% to 95.5%, while preserving sufficient (> 5 min) REM sleep for 95% of subjects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。