让机器用自然语言理解睡眠,实现从分析到对话的突破
SleepLM: Natural-Language Intelligence for Human Sleep
- 构建多层级睡眠描述生成流程,创建首个超10万小时的大规模睡眠文本数据集
- 统一预训练目标融合对比对齐、文本生成与信号重建,提升跨模态一致性
- 支持零样本任务泛化和语言引导事件定位,适合睡眠研究与智能健康应用
我们提出 SleepLM,一个面向人类睡眠的自然语言基础模型家族,实现睡眠状态的对齐、解析与自然语言交互。尽管睡眠至关重要,现有基于学习的睡眠分析系统仍局限于预定义标签空间(如特定睡眠阶段或事件),无法描述、查询或泛化至新睡眠现象。SleepLM 将自然语言与多模态多导睡眠图结合,实现睡眠生理的语言化表征。为此,我们设计了一种多层次睡眠描述生成流水线,构建了首个大规模睡眠-文本数据集,涵盖超过10万小时数据,来自1万余名个体。此外,我们提出一种统一预训练目标,整合对比对齐、文本生成与信号重建,更精准捕捉生理保真度与跨模态互动。在真实世界睡眠理解任务上的实验表明,SleepLM 在零样本与少样本学习、跨模态检索及睡眠描述生成方面均超越当前最优方法。更重要的是,SleepLM 展现出语言引导事件定位、针对性洞察生成及对未见任务的零样本泛化等能力。所有代码与数据将开源。
原文摘要 · Abstract (English)
We present SleepLM, a family of sleep-language foundation models that enable human sleep alignment, interpretation, and interaction with natural language. Despite the critical role of sleep, learning-based sleep analysis systems operate in closed label spaces (e.g., predefined stages or events) and fail to describe, query, or generalize to novel sleep phenomena. SleepLM bridges natural language and multimodal polysomnography, enabling language-grounded representations of sleep physiology. To support this alignment, we introduce a multilevel sleep caption generation pipeline that enables the curation of the first large-scale sleep-text dataset, comprising over 100K hours of data from more than 10,000 individuals. Furthermore, we present a unified pretraining objective that combines contrastive alignment, caption generation, and signal reconstruction to better capture physiological fidelity and cross-modal interactions. Extensive experiments on real-world sleep understanding tasks verify that SleepLM outperforms state-of-the-art in zero-shot and few-shot learning, cross-modal retrieval, and sleep captioning. Importantly, SleepLM also exhibits intriguing capabilities including language-guided event localization, targeted insight generation, and zero-shot generalization to unseen tasks. All code and data will be open-sourced.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。