arXiv:2608.12446cs.LGcs.AI2026-08

用多位专家标注数据构建更可靠的睡眠分期标签,提升自动分析准确性。

Personalized Scorer Modeling: A Learning-Based Framework for Deriving Robust Sleep Stage Labels from Multiple Experts

  • 基于混淆矩阵建模每位专家的判读习惯,融合多专家意见生成新标签。
  • 在两个公开数据集上达到85%以上准确率,优于单一专家或原标签。
  • 适合睡眠医学研究者和自动化睡眠分析系统开发者使用。

睡眠分期对睡眠障碍的诊断与管理至关重要,但现有自动分期研究多以单个参考脑电图(hypnogram)为基准,忽略了专家间评分差异。本研究探索利用多专家标注数据,通过集体行为构建更可靠的参考标签。采用公开数据集DOD-H和DOD-O,将脑电(C3-M2)与下颌肌电信号分割为30秒片段,每模态提取30个特征,共60维特征。提出一种基于学习的脑电图(LBH),通过机器学习模型生成的混淆矩阵,建模每位评分者在各睡眠阶段的误判模式;经列归一化后估计每个真实阶段的概率,并聚合多专家结果确定最终标签。在仅用脑电与脑电+肌电两种条件下,分别使用随机森林、支持向量机和多层感知机进行评估,对比原数据集脑电图(DH)与最优评分者脑电图(BSH)。LBH在所有设置中均表现更优,最佳结果为:在DOD-H上达86.07%准确率、85.46%精确率、85.29%F1值;在DOD-O上达86.04%准确率、85.21%精确率、84.70%F1值。结果表明,个性化评分者建模可有效提升参考标签可靠性,无需丢弃个体专家信息。

原文摘要 · Abstract (English)

Sleep stage classification is important for the diagnosis and management of sleep disorders, yet most automatic staging studies evaluate models against a single reference hypnogram despite known inter-scorer variability. This study investigates whether multi-scored datasets can be used to construct more reliable reference labels from the collective behavior of multiple experts. We use the publicly available DOD-H and DOD-O datasets. EEG (C3-M2) and chin EMG signals were segmented into 30-s epochs, and 30 features were extracted from each modality, yielding 60 features for EEG+EMG. We propose a learning-based hypnogram (LBH) that models the stage-specific behavior of each scorer using confusion matrices derived from machine-learning models. After column normalization, these matrices estimate the probability of each true sleep stage given each scorer's label; probabilities are aggregated across scorers to assign the final label for each epoch. LBH was evaluated with random forest, support vector machine, and multilayer perceptron classifiers under EEG-only and EEG+EMG settings, and compared with the dataset hypnogram (DH) and best-scorer hypnogram (BSH). LBH consistently improved overall performance. The best results were obtained with random forest and EEG+EMG, reaching 86.07% accuracy, 85.46% precision, and 85.29% F1-score on DOD-H, and 86.04% accuracy, 85.21% precision, and 84.70% F1-score on DOD-O. These findings suggest that personalized scorer modeling can improve reference hypnogram construction without discarding information from individual experts.

睡眠分期多专家融合机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。