用图像化方法让通用模型学会睡眼分期,更像医生诊断。
Introducing Multimodal Paradigm for Learning Sleep Staging PSG via General-Purpose Model
- 将脑电波信号转为可视图像,让大模型像看图一样分析睡眠阶段。
- 在三个数据集上达到领先准确率,无需专门训练睡眠数据。
- 模型学习了医生看图判读的思路,结果可解释性强,适合医疗场景。
睡眠分期对诊断睡眠障碍和评估神经健康至关重要。现有自动方法通常从复杂的多导睡眠图(PSG)信号中提取特征,并训练领域专用模型,往往缺乏直观性且需要大量专业数据。为克服这些局限,我们提出一种新型睡眠分期范式,利用大型多模态通用模型模拟临床诊断实践。具体地,我们将原始一维 PSG 时间序列转换为直观的二维波形图像,并微调多模态大模型从这些表示中学习。在三个公开数据集(ISRUC、MASS、SHHS)上的实验表明,该方法使通用模型在未接触过睡眠数据的情况下仍具备稳健的分期能力。此外,解释性分析显示,模型学会了模仿人类专家通过 PSG 图像进行睡眠分期的视觉诊断流程。所提方法在准确率与鲁棒性上持续优于现有最优基线,凸显其在医疗应用中的高效性与实用价值。信号到图像的代码管道及 PSG 图像数据集将公开发布。
原文摘要 · Abstract (English)
Sleep staging is essential for diagnosing sleep disorders and assessing neurological health. Existing automatic methods typically extract features from complex polysomnography (PSG) signals and train domain-specific models, which often lack intuitiveness and require large, specialized datasets. To overcome these limitations, we introduce a new paradigm for sleep staging that leverages large multimodal general-purpose models to emulate clinical diagnostic practices. Specifically, we convert raw one-dimensional PSG time-series into intuitive two-dimensional waveform images and then fine-tune a multimodal large model to learn from these representations. Experiments on three public datasets (ISRUC, MASS, SHHS) demonstrate that our approach enables general-purpose models, without prior exposure to sleep data, to acquire robust staging capabilities. Moreover, explanation analysis reveals our model learned to mimic the visual diagnostic workflow of human experts for sleep staging by PSG images. The proposed method consistently outperforms state-of-the-art baselines in accuracy and robustness, highlighting its efficiency and practical value for medical applications. The code for the signal-to-image pipeline and the PSG image dataset will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。