构建开放评估框架,让睡眠分期模型公平比拼。
SLEEPYLAND: trust begins with fair evaluation of automatic sleep staging models
- 提出SLEEPYLAND框架,含超22万小时真实睡眠数据
- 集成模型SOMNUS在24个数据集上平均准确率达87.2%
- 首次实现模型性能超越人类专家,且可预测标注分歧
尽管深度学习在自动睡眠分期方面取得进展,但临床应用受限于模型评估不公平、跨数据集泛化差、模型偏差及人工标注差异。本文提出SLEEPYLAND,一个开源睡眠分期评估框架,包含超过22万小时域内(ID)和8.4万小时域外(OOD)睡眠记录,覆盖广泛年龄、睡眠障碍类型与设备配置。我们发布了基于先进架构的预训练模型,并在单/多通道脑电图/眼电图设置下标准化评估。引入SOMNUS集成模型,通过软投票融合不同架构与通道组合,在24个数据集上宏平均F1得分介于68.7%至87.2%,94.9%情况下优于单一模型。值得注意的是,即使对比模型在域内训练,SOMNUS在域外仍表现更优。基于BSWR子集(N=6,633),量化了年龄、性别、AHI、PLMI相关的模型偏差,表明虽集成提升鲁棒性,但无一架构能始终降低偏差。在域外多标注数据集DOD-H和DOD-O上,SOMNUS超越最佳人类评分者(宏F1 85.2% vs 80.8%;80.2% vs 75.9%),更贴近评分共识(κ=0.89/0.85,ACS=0.95/0.94)。最后提出基于熵与模型间差异的集成不确定度指标,预测评分分歧的ROC AUC最高达0.828,提供数据驱动的人类不确定性代理。
原文摘要 · Abstract (English)
Despite advances in deep learning for automatic sleep staging, clinical adoption remains limited due to challenges in fair model evaluation, generalization across diverse datasets, model bias, and variability in human annotations. We present SLEEPYLAND, an open-source sleep staging evaluation framework designed to address these barriers. It includes more than 220'000 hours in-domain (ID) sleep recordings, and more than 84'000 hours out-of-domain (OOD) sleep recordings, spanning a broad range of ages, sleep-wake disorders, and hardware setups. We release pre-trained models based on high-performing SoA architectures and evaluate them under standardized conditions across single- and multi-channel EEG/EOG configurations. We introduce SOMNUS, an ensemble combining models across architectures and channel setups via soft voting. SOMNUS achieves robust performance across twenty-four different datasets, with macro-F1 scores between 68.7% and 87.2%, outperforming individual models in 94.9% of cases. Notably, SOMNUS surpasses previous SoA methods, even including cases where compared models were trained ID while SOMNUS treated the same data as OOD. Using a subset of the BSWR (N=6'633), we quantify model biases linked to age, gender, AHI, and PLMI, showing that while ensemble improves robustness, no model architecture consistently minimizes bias in performance and clinical markers estimation. In evaluations on OOD multi-annotated datasets (DOD-H, DOD-O), SOMNUS exceeds the best human scorer, i.e., MF1 85.2% vs 80.8% on DOD-H, and 80.2% vs 75.9% on DOD-O, better reproducing the scorer consensus than any individual expert (k = 0.89/0.85 and ACS = 0.95/0.94 for healthy/OSA cohorts). Finally, we introduce ensemble disagreement metrics - entropy and inter-model divergence based - predicting regions of scorer disagreement with ROC AUCs up to 0.828, offering a data-driven proxy for human uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。